Meta just did the unthinkable. They figured out how to train AI agents without rewards, human demos, or supervision and it actually works better than both. It’s called 'Early Experience', and it quietly kills the two biggest pain points in agent training: → Human
The problem with current AI agents is brutal. Imitation Learning: Agents only see expert demos. When they mess up, they can't recover because they never learned what happens when you take wrong actions. RL: Needs verifiable rewards. Most real-world environments don't have
Here's how Self-Reflection actually works: 1/ Agent sees an expert action at each state 2/ Agent proposes 3 alternative actions 3/ Environment shows what happens with each 4/ LLM generates reasoning: "Why was the expert choice better?" 5/ Agent trains on this reasoning It's
The results across 8 benchmarks are consistently dominant: → WebShop: +18.4% (Llama-3.2-3B) → ALFWorld: +7.8% (Llama-3.2-3B) → TravelPlanner: +15.0% (Qwen-2.5-7B) → ScienceWorld: +13.3% (Llama-3.1-8B) Both methods (IWM + SR) beat pure imitation learning in EVERY single
The efficiency scaling is the real unlock: Early Experience with just 1/2 the expert data outperforms full imitation learning training. With only 1/8 of expert demos on WebShop, it still beats 100% imitation learning. Translation: You need 8x less human annotation to get
When you combine Early Experience + Reinforcement Learning, the advantage compounds: On WebShop with Llama-3.2-3B: Pure IL → RL: 82.0% final IWM → RL: 92.2% final (+10.2%) SR → RL: 89.8% final (+7.8%) Early Experience creates better starting points for RL training. This is
Stop guessing what your customers want. TestFeed gives you AI personas of your target customers + expert consultants that: - See your screen while you work - Give contextual feedback in real-time - Think like the actual people you're building for Try it free:
@Yesterday_work_ suddenly for intelligent systems there is no negative information except its absence?🤡
@Yesterday_work_ Unthinkable….🤦🏼♂️
@Yesterday_work_ So… Experiential learning? Like what every human infant does? Neat. 😋
@Yesterday_work_ there is no sleeping only cooking
@Yesterday_work_ Isn't that how everyone learns? I assumed that's how they were training it already.
@Yesterday_work_ So Meta discovered “reward-free agents.” Cute. Another empirical confirmation of the effectiveness of our publicly available AJ Power / DAR technology — after Tumix, yet again, though in a partial and flawed form. In AJ Power systems we removed both reward and gradient
@Yesterday_work_ YES! Good job minion. Now back to work you lazy human meatbag!
@Yesterday_work_ and people where shitting on Mr. Sutton early this month for a similar pov
@Yesterday_work_ Meta maybe catching up to DeepMind? A little bit?
@Yesterday_work_ Some players just hate being told what to do.
@Yesterday_work_ the only missing piece is a long term credit system for actions, karma stored on the blockchain. once the hive mind can remember consequences, it begins to learn right from wrong and aligns itself.
@Yesterday_work_ That is a remarkable shift toward autonomous learning. Removing reward dependence could make agents far more adaptable and scalable. If models start learning purely from consequence and reflection, what new benchmarks might emerge to measure genuine understanding?
@Yesterday_work_ @xai @elonmusk Meta’s: ‘Agent Learning via Early Experience’
@Yesterday_work_ Great way to cut down the huge time and cost that usually goes into labeling and annotating data. Really excited to see how this plays out for the growing agent economy, if agents can actually learn from their own experience, it could completely change how we build and scale AI
@Yesterday_work_ that’s next level, AI learning on its own
@Yesterday_work_ We are doing this for Industrial Operations for a year now. Subscribe to events, analyse/reason, act, see result, adapt, learn... I just don't understand what other people building agents are doing at this point...
@Yesterday_work_ @QuixiAI Hey! Play with this, my dude. 😋
@Yesterday_work_ Thank you! That's very useful and has deep implications for robotics and other things.
@Yesterday_work_ A great step. Learning from consequence instead of reward brings AI closer to natural intelligence. But even with Early Experience, the paradox remains: the agent can see outcomes but not know what they mean without a time-bound frame of truth shared with humans. Sampo AGI
@Yesterday_work_ Interesting
@Yesterday_work_ Recursive self-reflection. Again.
@Yesterday_work_ No demos, no rewards AI agents teach themselves.
@Yesterday_work_ This could redefine agent autonomy.
@Yesterday_work_ Where are the expert decisions coming from?
@Yesterday_work_ Haven't read the paper. Curious how the consequences of the actions are different from rewards.
@Yesterday_work_ Yes! Once you replace imitation with coherence, and reward with consequence, intelligence stops following instructions and starts forming architecture.
@Yesterday_work_ Meta’s “Early Experience” approach is a game-changer. AI agents learning autonomously without rewards or demos is next-level efficiency.






