For friends of open source: imo the highest leverage thing you can do is help construct a high diversity of RL environments that help elicit LLM cognitive strategies. To build a gym of sorts. This is a highly parallelizable task, which favors a large community of collaborators.
@karpathy Perfect timing, we are just about to publish TextArena. A collection of 57 text-based games (30 in the first release) including single-player, two-player and multi-player games. We tried keeping the interface similar to OpenAI gym, made it very easy to add new games, and created
@karpathy Bytedance open sourced their RL framework for LLMs recently, plan to dive into it: https://github.com/volcengine/...
@karpathy With reasoning-gym we build procedurally generated & algorithmically verifiable datasets (envs) for RL training of reasoning models- already ~30 ready to use, goal is 100 for v1: https://github.com/open-though...
@karpathy I think this is where we should go after reproducing the initial results R1, expand the verifiable domains to medicine, law etc. I think the open source community is the perfect ground for this: https://github.com/huggingface...
@karpathy Karpathy either you leader such endeavor or nobody will (not saying you should, just stating a most likely fact)
@karpathy Here's one for agents doing scientific tasks https://github.com/Future-Hous...
@karpathy Data-centric ML > algorithm ML We organized an environment-centric RL workshop at NeurIPS 2021. Think beyond algorithmic design spaces; think how you can construct easier-to-RL but low-reward-hacking environments.
@karpathy We have a dozen envs that run 1M+ steps/second/core in pufferlib. All open source and several more in development by the community
@karpathy what you’re saying is you want to build an open ai gym?
@karpathy We created the first 2k software engineering environment in SWE-Gym: https://x.com/jiayi_pirate/sta... To scale this up, I really hope the open-source software/research community could find a way to standardize environment setup & CI/tests -- this could get us 100x+ SWE-Gym 🫡
@karpathy we did it 🫡 https://x.com/PrimeIntellect/s...
@karpathy So like a dog park but for robots? I love this
@karpathy reminds me of the olden days of openai gym (i know it's not related)
@karpathy virtual world projects like @hyperfy_io seem ideal for such; Hyperfy is a GPL licensed low-code collaborative real-time 3d web editor with AI integrations + semantic worlds: https://github.com/hyperfy-xyz... https://x.com/hyperfy_io/statu...
@karpathy RL is a good paradigm for decantrlized training as well
@karpathy From an Active Inference perspective, we are working on this along the lines of: Generative Research Teams: Active Inference Compositions For Research and Meta-Science https://zenodo.org/record/8164... An Active Inference Ontology for Decentralized Science: from Situated Sensemaking to
@karpathy I’m using RL to fine-tune this working model of the StoneyNakoda language. Data, models, and process all below. This is a RL linguistic “gym” as all native speakers are almost gone and there is little public data available that has been scraped. https://github.com/HarleyCoops...
@karpathy Have a simple LM-friendly text port of Craftax-Full here: https://github.com/JoshuaPurte...
@karpathy Sir. This is Bittensor.
@karpathy Any existing project you can point out to illustrate?
@karpathy Just make great benchmarks
@karpathy This is what @SentientAGI did with our werewolf "AGI-thon". We collected game transcripts from deductive agents playing werewolf and used SFT with filtered data to improve the model. Next step is to use RL for same https://github.com/sentient-ag...
@karpathy i would donate my pc’s night shift to a theorem proving gym, hope someone is working on this!
@karpathy Sounds like the AI party from the Torus Run https://thetorus.ai/
@karpathy Have you seen @onchaingaias and @henlokart
@karpathy Bittensor is trying and on the way to creating this, do you have any suggestion on how they can do it better?
@karpathy What is a cognitive strategy? Have we ever actually measured one? Or is RL just mimicking cognitive strategies? Can LLMs even model and store a cognitive strategy? All questions no one in this community has spent time asking before going all in. Mistake.
@karpathy Cool!
@karpathy So basically a crowdsourced intelligence gymnasium that can evolve to invent new exercises. Love it.
@karpathy Andrej, you might want to have a chat with @PaglieriDavide :-)
@karpathy OpenAI Gym 2.0?




