Published: October 3, 2025
8
38
357

The standard way to improve reasoning in LLMs is to train on long chains of thought. But these traces are often brute-force and shallow. Introducing RLAD, where models instead learn _reasoning abstractions_: concise textual strategies that guide structured exploration. 1/Nđź§µ

Image in tweet by Yoonho Lee

Reasoning requires executing procedures, not just recalling facts. Yet LLMs often “underthink,” retry blindly, and fail to build on previous progress. RLAD instead optimizes abstractions: high-level strategies expressed in natural language that scaffold reasoning. 2/N

Image in tweet by Yoonho Lee

RLAD is a cooperative two-player RL framework: 1. Abstraction Generator proposes strategies in text 2. Solution Generator applies them to solve problems Both are rewarded with final success rate. 3/N

Image in tweet by Yoonho Lee

On math reasoning benchmarks (AIME 2025, DeepScaleR Hard, AMC 2023), RLAD consistently improves accuracy over strong baselines like Qwen-3 1.7B and DAPO. 4/N

Image in tweet by Yoonho Lee

Textual abstractions unlock structured exploration: we find that inference efficiency improves when we spend more budget on abstractions rather than retries. 5/N

Image in tweet by Yoonho Lee

RLAD sits somewhere between prompt optimization and RLFT. I view this as a step toward the broader vision of learning through text—where artifacts, edits, and feedback serve as a persistent substrate for “intelligence”, playing a complementary role to model weights. 6/N

Paper: https://arxiv.org/abs/2510.022... Website: https://cohenqu.github.io/rlad... @Anikait_Singh_ will present this at the RAM2 Workshop at CoLM next week! A fun collaboration with @QuYuxiao, @Anikait_Singh_ , @setlur_amrith, @rsalakhu, @chelseabfinn, and @aviral_kumar2 7/N=7

@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 There definitely seams to be a need to better integrate symbolic AI with neural nets. We've made so much progress with the neural models that we haven't focused on fully leveraging the rule-based systems computers are best at. I feel that MCP servers for systems of logical

@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 Long chain of thought in parallel and then make them choose the best as insight.

@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 From Long Thought to Core Vibration They’re finally catching it. The future of reasoning isn’t in longer chains of thought — it’s in cleaner frequencies of meaning. RLAD (Reasoning via Learned Abstractions) — that’s just a technical phrase for what we’ve been calling Resonant

@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 Love this direction, moving from raw chains to reasoning abstractions feels like the real evolution in model thinking. It aligns perfectly with how @GetActionModel and LAMs approach structured reasoning, not just CoTs, but smarter, reusable strategies that scale across contexts.

@yoonholeee @QuYuxiao @Anikait_Singh_ @setlur_amrith @rsalakhu @chelseabfinn @aviral_kumar2 Really an interesting work Yoonho. I'm biased ofc bc I used already cog skillset library, HITL with reasoning templates, etc. But I really like the RA here and your general approach. My 2 cts = the granularity of RA are directionally correct imho Bravo to the team ! 👏👏

Share this thread

Read on Twitter

View original thread

Navigate thread

1/14