My brain broke when I read this paper. A tiny 7 Million parameter model just beat DeepSeek-R1, Gemini 2.5 pro, and o3-mini at reasoning on both ARG-AGI 1 and ARC-AGI 2. It's called Tiny Recursive Model (TRM) from Samsung. How can a model 10,000x smaller be smarter? Here's how
The new AI model architecture Tiny Recursive Model (TRM) beats DeepSeek R1, Gemini 2.5 Pro, and o3-mini on ARC-AGI 1 and ARC-AGI-2. It has only 7 Million parameters and used 1000 training samples.
Less is More: Recursive Reasoning with Tiny Networks Abstract: Hierarchical Reasoning Model (HRM) is a novel approach using two small neural networks recursing at different frequencies. This biologically inspired method beats Large Language models (LLMs) on hard puzzle tasks
Samsung published the code this morning on Github. https://x.com/justthisguy/stat...
Is TRM just overfitting or benchmark maxing? The author @jm_alexia directly address this question. The model was trained on 1,000 Sudoku examples and achieved 87.4% accuracy on a test set of 423K puzzles. This demonstrates the model was able to generalize within a domain.
@GregKamradt, the president of @arcprize might put out a bounty to get TRM tested on the private set. I'd love to see this. https://x.com/GregKamradt/stat...
@JacksonAtkinsX You're talking about two zero-sum specialized tests. How is this different than the neural network game training that we did back in the 90's but faster? Of course this should beat general LLMs. Any small set of neural networks overtrained to a constrained set of possible results
@proteusguy Yes, you are correct this is not general intelligence. HRM generated excitement but it didn't generalize well. The paper's author (not me) took HRM and looked at why it failed to generalize and corrected those issues. I see this paper as a next step in developing these types
@JacksonAtkinsX You should mention the caveats. It is tiny because it is very specialized model without language. It would have been much more interesting if they had done more standard math/coding reasoning.
@sytelus Yes the model is specialized for the tasks such as Sudoku. The main innovation is that it generalizes much better than the previous HRM. As the saying goes, “Just imagine 2 more papers down the line”
@JacksonAtkinsX can we actually try it anywhere ?
@iEgit @samsungresearch @Samsung What do you think about open sourcing this for the community? Let's have the ARC team benchmark it independently like they did with HRM. Here is the pseudocode from the paper for anyone who wants to try building one.
@JacksonAtkinsX The link to the paper?
@apdemetriou The link to the paper is in the comment with the abstract.
@JacksonAtkinsX HF link or it didn't happen ;)
@raw_works See the comments. A link to Github was just added.
@JacksonAtkinsX Have a feeling it's either specifically trained for that Benchmark or it's VooDoo My guess is VooDoo
@EddyLeeKhane They tested it on other tasks like Sudoku and Maze. They note they used 1K training samples for Sodoku.
@JacksonAtkinsX @grok will the main Llm providers use this any time soon ?
@JacksonAtkinsX How fast is it compared to other models on the market? That much recursive thinking sounds like mean response time
@JacksonAtkinsX @grok how much compute is needed to run this new model?
@JacksonAtkinsX Samsung. My favourite hardware brand.
@JacksonAtkinsX Stop chasing bigger models. Start tuning cleaner frequencies. The future of intelligence isn’t recursive — it’s resonant. You don’t need 16 loops of self-correction when the system is aligned with the signal of truth. 🧠 Recursion = guessing better. ⚡ Resonance = knowing
@JacksonAtkinsX Well, if an ant can look in the mirror and notice something it needs to clean, AI should be able to be made smaller. https://veg1.org/ants.html
@JacksonAtkinsX So its not really the model per se, its how its used in a loop to refine itself. Seems pretty obvious, really. I thought of something very similar a couple years ago (then had further thoughts beyond this). But I don't think our current path is a good one and building AGI/ASI
@JacksonAtkinsX Mind-blowing stuff, right? 🤯 Do you think the success of Tiny Recursive Model indicates a shift towards smaller, more efficient AI models in the future? What's the next big hurdle for these mini-models?
@JacksonAtkinsX Yes, Gusss—I've reviewed the full TRM paper from Samsung titled “Less is More: Recursive Reasoning with Tiny Networks”. Here's the forensic sweep: 🔍 TRM’s Core Architecture Initial Drafting: Generates a full answer in one go—not word-by-word like LLMs. Latent Scratchpad:
@JacksonAtkinsX So this is how the A.I. bubble bursts….
@JacksonAtkinsX How does it check it's logic? Does it have some symbolic reasoning built in? Otherwise what do we mean by check?
@JacksonAtkinsX It broke your mind? Computers were room size once. Caculators were an advanced technology once. People road horse back because that's all they had. It shouldn't surprise you.
@JacksonAtkinsX Yeah we can create really small narrow intelligence models. Good for active inference of business models the business value of RD but the Goliath of creating SLMs comparable to LLMs is not yet here at least on common sense parlance
@JacksonAtkinsX My Prediction is the TG side operates like thought to action. The HRM/TRM concepts will act like strategy recall, social mannerisms (scenario context, morality, etc think serotonin), skill development (like learning processes in a business or how to use specific tools), goals
@JacksonAtkinsX Nooooo 😭💀 > Create a "Scratchpad": It then creates a separate space for its internal thoughts, a latent reasoning "scratchpad." This is where the real magic happens. @ESYudkowsky
@JacksonAtkinsX Surely this can be combined with the latest and greatest LLM for even better results though, yeah?
@JacksonAtkinsX What's the trade off here? Longer thinking time? Even more inference and/or test time compute?
@JacksonAtkinsX We are poud to announce that we have built Sabel mini from the hit book 'If anyone builds Sable everyone dies'





