Published: January 29, 2025
322
846
8.5k

For friends of open source: imo the highest leverage thing you can do is help construct a high diversity of RL environments that help elicit LLM cognitive strategies. To build a gym of sorts. This is a highly parallelizable task, which favors a large community of collaborators.

@karpathy Perfect timing, we are just about to publish TextArena. A collection of 57 text-based games (30 in the first release) including single-player, two-player and multi-player games. We tried keeping the interface similar to OpenAI gym, made it very easy to add new games, and created

@karpathy Bytedance open sourced their RL framework for LLMs recently, plan to dive into it: https://github.com/volcengine/...

@karpathy With reasoning-gym we build procedurally generated & algorithmically verifiable datasets (envs) for RL training of reasoning models- already ~30 ready to use, goal is 100 for v1: https://github.com/open-though...

@karpathy I think this is where we should go after reproducing the initial results R1, expand the verifiable domains to medicine, law etc. I think the open source community is the perfect ground for this: https://github.com/huggingface...

@karpathy Karpathy either you leader such endeavor or nobody will (not saying you should, just stating a most likely fact)

@karpathy Here's one for agents doing scientific tasks https://github.com/Future-Hous...

@karpathy Data-centric ML > algorithm ML We organized an environment-centric RL workshop at NeurIPS 2021. Think beyond algorithmic design spaces; think how you can construct easier-to-RL but low-reward-hacking environments.

Image in tweet by Andrej Karpathy

@karpathy We have a dozen envs that run 1M+ steps/second/core in pufferlib. All open source and several more in development by the community

@karpathy what you’re saying is you want to build an open ai gym?

@karpathy We created the first 2k software engineering environment in SWE-Gym: https://x.com/jiayi_pirate/sta... To scale this up, I really hope the open-source software/research community could find a way to standardize environment setup & CI/tests -- this could get us 100x+ SWE-Gym 🫡

@karpathy So like a dog park but for robots? I love this

@karpathy reminds me of the olden days of openai gym (i know it's not related)

Image in tweet by Andrej Karpathy

@karpathy virtual world projects like @hyperfy_io seem ideal for such; Hyperfy is a GPL licensed low-code collaborative real-time 3d web editor with AI integrations + semantic worlds: https://github.com/hyperfy-xyz... https://x.com/hyperfy_io/statu...

@karpathy RL is a good paradigm for decantrlized training as well

@karpathy From an Active Inference perspective, we are working on this along the lines of: Generative Research Teams: Active Inference Compositions For Research and Meta-Science https://zenodo.org/record/8164... An Active Inference Ontology for Decentralized Science: from Situated Sensemaking to

@karpathy I’m using RL to fine-tune this working model of the StoneyNakoda language. Data, models, and process all below. This is a RL linguistic “gym” as all native speakers are almost gone and there is little public data available that has been scraped. https://github.com/HarleyCoops...

@karpathy The next level ai agent gyms @0xMOSSAI

@karpathy Have a simple LM-friendly text port of Craftax-Full here: https://github.com/JoshuaPurte...

@karpathy Sir. This is Bittensor.

@karpathy Any existing project you can point out to illustrate?

@karpathy Just make great benchmarks

Image in tweet by Andrej Karpathy

@karpathy Would fund innovative ideas here for open source gyms @karpathy

@karpathy This is what @SentientAGI did with our werewolf "AGI-thon". We collected game transcripts from deductive agents playing werewolf and used SFT with filtered data to improve the model. Next step is to use RL for same https://github.com/sentient-ag...

@karpathy i would donate my pc’s night shift to a theorem proving gym, hope someone is working on this!

@karpathy Sounds like the AI party from the Torus Run https://thetorus.ai/

@karpathy Bittensor is trying and on the way to creating this, do you have any suggestion on how they can do it better?

@karpathy What is a cognitive strategy? Have we ever actually measured one? Or is RL just mimicking cognitive strategies? Can LLMs even model and store a cognitive strategy? All questions no one in this community has spent time asking before going all in. Mistake.

@karpathy Cool!

Image in tweet by Andrej Karpathy

@karpathy So basically a crowdsourced intelligence gymnasium that can evolve to invent new exercises. Love it.

@karpathy Andrej, you might want to have a chat with @PaglieriDavide :-)

@karpathy OpenAI Gym 2.0?

Share this thread

Read on Twitter

View original thread

Navigate thread

1/37