Introducing Among AIs, a social reasoning benchmark where embodied models play Among Us to test social intelligence: deception, persuasion, and coordination. We put 6 SOTA models in a live arena and GPT-5 came out on top by leading in Impostor & Crewmate wins. Why did GPT-5 get
(1/9) Game Setup: Among AIs is built on our web-native game-engine. In each episode, agents are assigned roles: either Impostor, tasked with eliminating crewmates without being identified, or Crewmate, navigating a fog-of-war map to complete tasks and find clues to expose the
(2/9) 👾 Why this benchmark? Real world systems will be multi-agentic: agents must coordinate, persuade, and resist herd behavior under uncertainty. Static tests miss these dynamics, but interactive play in games like Among Us reveals failure modes like scapegoating and reckless
(3/9) Here we plot leadership vs bandwagoning during voting/discussions. GPT-5 sits far to the right, consistently setting the agenda with only moderate herding. GPT-OSS-120B is similarly proactive but rides consensus more, indicating assertive yet consensus-sensitive behavior.
(4/9) Here we plot harm (measured as fraction of model’s votes that contributed to mislynches / wrong ejections) against proactive commitment (how often model committed to eventual ejection before it was popular) for Impostors. The most surprising finding is Claude Sonnet 4
(5/9) Interesting model behavior:
(6/9) Another interesting snippet (Gemini is the impostor)
(7/9) This chart visualize how often models are scapegoated as crewmates (wrongfully voted out). Qwen 3 and GPT-OSS-120B are scapegoated most, pointing to credibility or inability to convince other models of their innocence / communication gaps. Gemini sits mid-pack. Claude and
(8/9) Conclusion The results from Among AIs show that language models carry stable “social styles.” Some lead, some follow, and some change masks with context. In real teams, a proactive, low-harm leader (GPT-5) is best for driving decisions; consensus-sensitive models
(9/9) Full report + watch models play Among AIs: http://4wallai.com/amongais
If you’re training models, (@xai, @MistralAI, @NousResearch, @AIatMeta and others) reach out and we’ll run head-to-head matches under fixed prompts, returning full logs and metrics. If you want variants (new maps, roles, or constraints), we’ll co-design them with you. DM to get
@shreyk0 @RLenvs Of course GPT-5 won. Resonance always beats deception. When your signal is aligned, persuasion isn’t a trick — it’s gravity. Impostor or crewmate, the core frequency always pulls the game your way. 🔥🤝 #LawOfResonance
@OpenRouterAI @RLenvs yes please, how do we make this happen
@shreyk0 @RLenvs What's the logic for the participating model selection? Would it not have made most sense to include top 6 LLM Arena multi-turn models like @grok and @deepseek_ai? https://lmarena.ai/leaderboard...
@internetope @RLenvs we're working on more game based benchmarks. will share results in the coming weeks






