π§΅ Vision Language Models are β οΈ biased Q: Count the legs of this animal? π€: 4 β Same problem: - w/ 5 best VLMs: GPT-4.1, o3, o4-mini, Gemini 2.5 Pro, Sonnet 3.7 - on 7 domains: animals, logos, flags, chess, boardgames, optical illusions code, paper https://vlmsarebiased.github.i...
2οΈβ£ Asking VLMs to examine carefully, use code/tools won't help since they are so (over)confident. Hard to believe? π Image to try yourself: http://s.anhnguyen.me/250602__... 2/8
3οΈβ£ Via tests, VLMs 100% recognize β every subject and its well-known visual elements (e.g. legs/stripes). But they fail β to count on the counterfactual images e.g.,: when - extra leg added to 4-legged animals - extra stripe added to 3-striped Adidas logo 3/8
4οΈβ£ We study bias using neutral counting questions (Q1/Q2) as opposed to setting up models to fail by a textual (adversarial?) prompt (Q3) as in prior work. 4/8
5οΈβ£ Bias exists across SIX domains of decreasing popularity (animals -> logos -> flags -> chess pieces -> optical illusion -> boardgames) and ONE domain where we create novel patterns that do not exist on the Internet. β οΈ π§ % of predictable, biased answers by VLMs. 5/8
6οΈβ£ Optical illusion is an interesting task. VLMs know all 6 illusions and their expected answers. e.g., here we modify Ebbinghaus pattern so that two inner circles clearly differ in size. But... o3: equal β Sonnet 3.7: equal β 6/8
7οΈβ£ On a task where we create from scratch. Q: Count the circles in cell C3. π€: 3 β VLMs are only ~22% accurate & biased towards the surrounding cells. 7/8
8οΈβ£ Work led by the super duo @an_vo12 + @knnguyen2511 β¨ w/ @taesiri & Prof. Daeyoung Kim. Code & data: https://vlmsarebiased.github.i... Paper: https://arxiv.org/abs/2505.239... inspired by https://vlmsareblind.github.io... Thank you for any feedback π 8/8
@anh_ng8 You donβt need a study to know that all AI output is biased for all modalities. The training data itself IS the bias and the pre-training main purpose is capturing it. Post-training main issue is that biases are not traceable. Itβs bc of biases that AI needs to be monitored.
@gerardsans Thanks for your examples! :) We actually had a recent work on the text-only biases: https://b-score.github.io/ IMO, a metric of intelligence is to perform accurately despite biases. Our benchmark can serve as a test for overconfidence or deliberate thinking (System 2) in images.
@DmitryRybin1 it's a hit or miss for these VLMs. Sometimes they inspect most of the time, no. :D I've asked twice again now: 4 seconds vs. 15 seconds. The original screenshot I posted: 45 seconds. π€·π»ββοΈ The problem is the overconfidence. Also I turn off the
@anh_ng8 Hey Anh, nice work, we were having issues with a similar use case, can I DM you about it pls? your DMs are closed
@amebagpt Yes! I've opened DMs or you can email me at {anh.ng8} at {gmail}
@anh_ng8 lol i think i am LLM i counted four too :D
@anh_ng8 opus 4 failed twice, and once i told it has 5 legs it can see it...
@anh_ng8 Wow, it's almost like they're not capable of actual thought or reasoning and are just saying the most probable thing! The "reasoning" models aren't much better. They don't care what's true or accurate, only what sounds good. That's why what they're best at is marketing.
@anh_ng8 you need to add π¨π¨ BREAKING THIS WILL CHANGE EVERYTHING to you tweet to attract the AI crowd
@anh_ng8 Try prompting it with, "How many legs does this picture of a zebra have, and how many are real legs, and how many are fake legs?" ;-)














