Published: September 23, 2025
5
5
22

Transformers are broken. Today, Manifest AI is releasing Power Retention, an open-source architecture to replace them. More below 🧵:

Transformers are incredibly powerful, but this power comes at a cost. Transformers memorize everything that they encounter, and at long context, this becomes a bottleneck. Rich understanding *or* long context: transformers can only have one.

Manifest's Power Retention enables LLMs to reason efficiently across millions of tokens. Applications that were previously out of reach due to cost of context become feasible. Assistants without session resets, days-long reasoning traces, full software engineering workflows...

Trying out Power Retention is as easy as `pip install retention`. We’re also releasing: - PowerCoder 3B: a super-fast, long-context code autocompletion model showcasing our tech - Vidrial: our framework for clean, high-performance CUDA kernels

For a deeper dive, our paper "Scaling Context Requires Rethinking Attention" is available on ArXiv: https://arxiv.org/pdf/2507.042...

Power Retention represents a step change for the field. We've done nothing but exploit scale for the past five years; it's time to rethink the fundamentals. No more hacks, patches, or RAG. The long context era is about to begin. Are you ready?

Discover more details, access the code, benchmarks, and read the full paper here: https://manifestai.com/article...

@jacobmbuckman MORE POWER. MORE RETENTION.

@jacobmbuckman Agents that can stay on task for weeks at a time! Now you have my attention. Want to be a guest on my podcast? this is the vibe: https://www.youtube.com/watch?...

@IanTimotheos Sure! DM me

@jacobmbuckman congrats on the release! agreed that language is overly dominated by localized context for training, we’ve been kicking the tires on hnets for similar reasons. excited to see how you build on this!

@jacobmbuckman power retention, huh? curious about runtimes?

@perscentsuply At 64k context, ~10x faster training and ~100x faster inference. Even larger contexts come with correspondingly larger gains.

Share this thread

Read on Twitter

View original thread

Navigate thread

1/13