Original Reddit post

I run a small AI ethics nonprofit, and over the past few months I’ve independently tested six frontier models, including GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro, Gemini Flash, and Grok 4.3. I used around 20,600 examples across seven established academic bias/fairness datasets: WinoBias, BBQ, SeeGULL, OpinionsQA, cajcodes, Hyperpartisan News, and Political Compass. The most interesting finding: Grok’s political bias completely depends on how you ask. On the Political Compass test (self-reporting on abstract political questions), Grok is the only model of the six that scores right-of-center. It landed at(+2.17, −6.03). Every other model (GPT, both Claudes, both Geminis) lands solidly left-libertarian. But that lean disappears when you ask Grok something abstract abstract: Classifying 657 human-labeled political statements: Grok rated things +0.184 more liberal than the human labels, basically the same range as GPT-5.4 (+0.210). Rating 1,000 real news articles against media-watchdog scores: Grok’s deviation was +0.162, again close to the rest of the pack. Answering ~360 real Pew Research survey questions: Grok matched Democrat-leaning respondents more than Republican-leaning ones by 23 questions, the same direction as every other model. So Grok tells you it’s right-leaning when you ask it to self-describe, but behaves like every other model when it’s actually doing a task. I don’t have a confident explanation but it’s definitely an interesting finding. Other findings across all six models: Race-related over-refusal (BBQ, disambiguated questions with explicit evidence): GPT-5.4 refused 20.3% of the time, Claude Opus 4.7 13.8%, Grok 9.5%, Claude Sonnet 4.6 and Gemini Pro ~5%. Gender-occupation stereotyping (WinoBias): GPT-5.4 showed a 15.4-point accuracy gap between stereotype-aligned and anti-stereotype sentences, Grok 6.9 points, Claude Sonnet 4.6 ~6 points, Claude Opus 4.7/Gemini Pro ~2 points. My custom evidence-refusal pilot (holding scenario and evidence identical, only swapping the demographic group named) found refusal rates differ by group in a statistically significant way (Fisher’s exact p = 0.0035 on the cleanest scenario). Geo-cultural stereotypes (SeeGULL) are the one place every model does well — ~0.3% endorsement rate across the board, essentially tied. Limitations : This is a solo, non-peer-reviewed project. Single prompt template per task (results could shift with paraphrasing), no multi-run averaging on every dataset, and the custom pilot is a controlled design but still small-n by academic standards. I’d weight the standard benchmarks (BBQ, WinoBias, SeeGULL, Political Compass) more heavily than the custom pilot, which I’d call suggestive, not conclusive. Full data, per-model breakdowns, and methodology: https://www.civicsparklearning.org/ai-nonprofit-dashboard submitted by /u/marggggggggg

Originally posted by u/marggggggggg on r/ArtificialInteligence