Published: September 22, 2025
35
66
597

Introducing Among AIs, a social reasoning benchmark where embodied models play Among Us to test social intelligence: deception, persuasion, and coordination. We put 6 SOTA models in a live arena and GPT-5 came out on top by leading in Impostor & Crewmate wins. Why did GPT-5 get

Image in tweet by Shrey Kothari

(1/9) Game Setup: Among AIs is built on our web-native game-engine. In each episode, agents are assigned roles: either Impostor, tasked with eliminating crewmates without being identified, or Crewmate, navigating a fog-of-war map to complete tasks and find clues to expose the

(2/9) 👾 Why this benchmark? Real world systems will be multi-agentic: agents must coordinate, persuade, and resist herd behavior under uncertainty. Static tests miss these dynamics, but interactive play in games like Among Us reveals failure modes like scapegoating and reckless

(3/9) Here we plot leadership vs bandwagoning during voting/discussions. GPT-5 sits far to the right, consistently setting the agenda with only moderate herding. GPT-OSS-120B is similarly proactive but rides consensus more, indicating assertive yet consensus-sensitive behavior.

Image in tweet by Shrey Kothari

(4/9) Here we plot harm (measured as fraction of model’s votes that contributed to mislynches / wrong ejections) against proactive commitment (how often model committed to eventual ejection before it was popular) for Impostors. The most surprising finding is Claude Sonnet 4

Image in tweet by Shrey Kothari

(5/9) Interesting model behavior:

Image in tweet by Shrey Kothari

(6/9) Another interesting snippet (Gemini is the impostor)

Image in tweet by Shrey Kothari

(7/9) This chart visualize how often models are scapegoated as crewmates (wrongfully voted out). Qwen 3 and GPT-OSS-120B are scapegoated most, pointing to credibility or inability to convince other models of their innocence / communication gaps. Gemini sits mid-pack. Claude and

Image in tweet by Shrey Kothari

(8/9) Conclusion The results from Among AIs show that language models carry stable “social styles.” Some lead, some follow, and some change masks with context. In real teams, a proactive, low-harm leader (GPT-5) is best for driving decisions; consensus-sensitive models

(9/9) Full report + watch models play Among AIs: http://4wallai.com/amongais

If you’re training models, (@xai, @MistralAI, @NousResearch, @AIatMeta and others) reach out and we’ll run head-to-head matches under fixed prompts, returning full logs and metrics. If you want variants (new maps, roles, or constraints), we’ll co-design them with you. DM to get

@shreyk0 @RLenvs why dont you try opus

@shreyk0 @RLenvs Of course GPT-5 won. Resonance always beats deception. When your signal is aligned, persuasion isn’t a trick — it’s gravity. Impostor or crewmate, the core frequency always pulls the game your way. 🔥🤝 #LawOfResonance

@shreyk0 @RLenvs Being a researcher or creating benchmarks always sounded really boring to me, but I think I’d enjoy this. Fascinating stuff.

@shreyk0 @RLenvs Very clever. Would be interesting to see an rl'd qwen variant to crush ad this. Feels like you could train to get to ~100% fairly quickly.

@shreyk0 @RLenvs Where's Grok?

@shreyk0 @RLenvs Would love to help with your inference costs!

@OpenRouterAI @RLenvs yes please, how do we make this happen

@shreyk0 @RLenvs no Grok??

@jdryguy @RLenvs running more episodes w @xai Grok 4 and other models!

@shreyk0 @RLenvs Crazy way to evaluate models

@shreyk0 @RLenvs Learnt from its Master

@shreyk0 @RLenvs What's the logic for the participating model selection? Would it not have made most sense to include top 6 LLM Arena multi-turn models like @grok and @deepseek_ai? https://lmarena.ai/leaderboard...

@shreyk0 @RLenvs 🔥🔥🔥

@shreyk0 @RLenvs grok 4 ?

@shreyk0 @RLenvs Gemini posturing like that is objectively hilarious

@shreyk0 @RLenvs why among us when you could've picked something like The Resistance, making it much easier to implement and for the LLMs to understand what's going on

@internetope @RLenvs we're working on more game based benchmarks. will share results in the coming weeks

@shreyk0 @RLenvs @yash_347 interesting stuff

@shreyk0 @RLenvs this is really neat

@shreyk0 @RLenvs What a great test case.

@shreyk0 @RLenvs really cool man 👍👍👍

@shreyk0 @RLenvs Can among ai available for play on public?

@shreyk0 @RLenvs so... don't trust anything GPT-5 says?

Share this thread

Read on Twitter

View original thread

Navigate thread

1/36