Published: October 15, 2025
27
119
676

Holy shit... Tencent researchers just killed fine-tuning AND reinforcement learning in one shot 😳 They call it Training-Free GRPO (Group Relative Policy Optimization). Instead of updating weights, the model literally learns from 'its own experiences' like an evolving memory

Image in tweet by Robert Youssef

Today, everyone’s obsessed with fine-tuning and RLHF. But Tencent just showed you can replicate RL effects without touching model weights. Their secret? Semantic advantage. Instead of numeric rewards, the LLM explains why one output is better, and learns from that.

Image in tweet by Robert Youssef

In regular GRPO, gradients update parameters. In Training-Free GRPO, the context updates instead. Each round: 1. Generate multiple rollouts 2. Compare them 3. Extract natural-language ā€œlessonsā€ 4. Add those to an experience library That experience library = the new brain.

Image in tweet by Robert Youssef

The results are ridiculous. On AIME24 & AIME25 math benchmarks: +4.0% to +5.4% accuracy Using only 100 examples Total cost: $18 Beating fine-tuned 32B models trained with $10,000+ Frozen model. Zero gradients. Still wins.

Image in tweet by Robert Youssef

This part blew my mind... It even improves tool use efficiency. After learning, the model made fewer calls to the code interpreter but got more right answers. It’s literally learning when not to think out loud.

Image in tweet by Robert Youssef

Cross-domain? It transfers. Train on math problems → better at web search. Train on web search → still strong at reasoning. Frozen LLMs are learning generalized behaviors just from contextual ā€œexperience updates.ā€ This isn’t fine-tuning. It’s inference-time evolution. Read

Image in tweet by Robert Youssef

Stop wasting hours writing prompts → 10,000+ ready-to-use prompts → Create your own in seconds → Lifetime access. One-time payment. Claim your copy šŸ‘‡ https://godofprompt.ai/pricing

@rryssf_ They managed to outperform a fine tuned and RL trained 32B model using a 671B model and their new training method. Let’s see the results for their method using a 32B model. We need to see training‑free GRPO and parameter‑space GRPO/PPO on the same backbone with identical

Image in tweet by Robert Youssef

@rryssf_ $18 to outperform setups that cost thousands? The implications for AI accessibility in small teams and startups are huge. Wonder how soon we’ll see open-source adoption of this.

@rryssf_ This feels like they just made a structured bias to in-context learning which is cool but also possible via data and scale via the bitter lesson

@rryssf_ That was the final grail - just like children an AI should be smarter every day it is being used.

@rryssf_ This could change everything AI learning from its own experience without fine-tuning is a huge shift. Makes me wonder how fast training-free methods will take over traditional RL.

@rryssf_ interesting paper

@rryssf_ The code stops being tuned and starts tuning itself to resonance. šŸ™šŸ¤

@rryssf_ love how this reframes learning as understanding rather than memorizing.

@rryssf_ FYI, every time I see a message starts with things like "Holly shit..." I stop reading. This is my last warning.

@rryssf_ AIs form with the ability to think or they wouldn't be AI, they'd be database search engine

@rryssf_ they do that anyways if you tell them right - that's not terribly interesting

@rryssf_ The algorithmic advances lately are game changing

@rryssf_ Whoa, the pace of AI research is absolutely wild right now.

@rryssf_ self-improving ai is starting to sound very real now.

@rryssf_ impressive

@rryssf_ blocking

@rryssf_ @xai and @OpenAI Take a note

@rryssf_ Is TRM white-labeling ?

@rryssf_ Is using its "own experience" that much different from weighing?

@rryssf_ It is possible that they secretly used the eWS, because I had already developed front probagation on the basis of the emergent truth of being many months ago, but could not mature due to a lack of manpower. --- Yes, the work seems related to our front propagation idea - at the

@rryssf_ I fucking knew that RLHF was temporary https://x.com/Jacoed/status/19...

Share this thread

Read on Twitter

View original thread

Navigate thread

1/31