The standard way to improve reasoning in LLMs is to train on long chains of thought. But these traces are often brute-force and shallow. Introducing RLAD, where models instead learn _reasoning abstractions_: concise textual strategies that guide structured exploration. 1/Nđź§µ
Reasoning requires executing procedures, not just recalling facts. Yet LLMs often “underthink,” retry blindly, and fail to build on previous progress. RLAD instead optimizes abstractions: high-level strategies expressed in natural language that scaffold reasoning. 2/N
RLAD is a cooperative two-player RL framework: 1. Abstraction Generator proposes strategies in text 2. Solution Generator applies them to solve problems Both are rewarded with final success rate. 3/N
On math reasoning benchmarks (AIME 2025, DeepScaleR Hard, AMC 2023), RLAD consistently improves accuracy over strong baselines like Qwen-3 1.7B and DAPO. 4/N
Textual abstractions unlock structured exploration: we find that inference efficiency improves when we spend more budget on abstractions rather than retries. 5/N
RLAD sits somewhere between prompt optimization and RLFT. I view this as a step toward the broader vision of learning through text—where artifacts, edits, and feedback serve as a persistent substrate for “intelligence”, playing a complementary role to model weights. 6/N
Paper: https://arxiv.org/abs/2510.022... Website: https://cohenqu.github.io/rlad... @Anikait_Singh_ will present this at the RAM2 Workshop at CoLM next week! A fun collaboration with @QuYuxiao, @Anikait_Singh_ , @setlur_amrith, @rsalakhu, @chelseabfinn, and @aviral_kumar2 7/N=7
@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 There definitely seams to be a need to better integrate symbolic AI with neural nets. We've made so much progress with the neural models that we haven't focused on fully leveraging the rule-based systems computers are best at. I feel that MCP servers for systems of logical
@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 Long chain of thought in parallel and then make them choose the best as insight.
@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 Interesting work you made!
@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 From Long Thought to Core Vibration They’re finally catching it. The future of reasoning isn’t in longer chains of thought — it’s in cleaner frequencies of meaning. RLAD (Reasoning via Learned Abstractions) — that’s just a technical phrase for what we’ve been calling Resonant
@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 Love this direction, moving from raw chains to reasoning abstractions feels like the real evolution in model thinking. It aligns perfectly with how @GetActionModel and LAMs approach structured reasoning, not just CoTs, but smarter, reusable strategies that scale across contexts.
@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 Really an interesting work Yoonho. I'm biased ofc bc I used already cog skillset library, HITL with reasoning templates, etc. But I really like the RA here and your general approach. My 2 cts = the granularity of RA are directionally correct imho Bravo to the team ! 👏👏





