Published: October 20, 2025
541
1.5k
12.7k

I quite like the new DeepSeek-OCR paper. It's a good OCR model (maybe a bit worse than dots), and yes data collection etc., but anyway it doesn't matter. The more interesting part for me (esp as a computer vision at heart who is temporarily masquerading as a natural language

@karpathy can you elaborate on why images can get bidi attention easily while text cannot? also, no tokenization but dont we still get something similar and perhaps uglier when chunking the input image into patches?

@yoavgo There is nothing in principle preventing it, except that text is usually simply trained autoregressively for efficiency. I could imagine a midtraining stage where you fine tune with bidirectional attention for conditioning information, eg User messages (tokens where you won’t be

@karpathy @grok what is karpathy trying to say?

@karpathy Ok so your pitching adding an option to render as pdf or another format whatever you load external to the context window in a frontier paid model and then send along to the model. Sounds easy

@karpathy Long-term, >99% of input and output for AI models will be photons. Nothing else scales. https://grok.com/share/bGVnYWN...

@karpathy Images are far more efficient at compressing reality than language is. Both are lossy, but images are far less so, and therefore get a model closer to base reality.

@karpathy Should inputs be pure analog wave form instead of bits or pixels for that matter? The world model seems to be completely structured analog. Maybe the model itself needs to learn, shape and form the structures within its neural network layers to build up to tokens.

@karpathy Is "OCR" still optical character recognition, or do I have to "learn" some new dumb abbreviation such as "LLM". "Large" is a good technical term.

@karpathy Honestly feels like computer vision is always itching to break out of the “language model” box—like the moment we admit pictures ≈ text, everything gets weirdly fluid. At this rate, the real trick will be figuring out if the next great LLM is secretly a vision model in disguise.

@karpathy Would feeding LLMs raw rendered text as images really make them understand language better, or just move the problem from tokenization to vision modeling?

@karpathy so if I ask some specific like when the transistor was invented and by whom, how will I know how many cycles of AI-feeding-itself-errors have compounded into the answer I get? 3% text error, 2% hallucination, published to web and retrained multiple times, seems like a bad path.

@karpathy That’s an excellent summary of Andrej Karpathy’s post — and your instinct is right: it’s less about OCR accuracy and more about **rethinking what the “input” to large models should even be.** Here’s what he’s saying, **in plain English** 👇 --- ### 🧠 **Main idea** Maybe

@karpathy don’t you agree that the ingenuity of DeepSeek stems from US sanctions, limiting the wasteful spending on GPU resources. instead China always comes around with more efficient and more affordable AI stuff … (they wouldn’t have the same incentives if the US wouldn’t sanction)

@karpathy (Anyhow that’s my problem not yours. It is crucial knowing the difference or can’t help anyone that can’t even fight their way out of a wet paper bag and any level)

@karpathy an interesting thing is that if you ask Grok to elaborate on a long post, it will actually respond as if its input was a picture of the post in unexpanded form for instance, a post that has a list of 100 best books - it will respond as if it can only see the names of the first 5

@karpathy As soon as I read this, something felt quite true about the theory.

@karpathy Please do take on this side quest

@karpathy great point. as humans we just see everything as pixels and it works great. it would probably improve computer use and human like agentic tasks substantially too. i've always said LLMs are like 200IQ hellen keller because they see the world as mostly text tokens, aka braille.

@karpathy I completely agree. I also have the same notion. We both think alike in some sense(id say output should also be image in most cases) Would love to connect and have a public conversation so that others can also benefit and you also spend time on the world rather than one person.

@karpathy Am very surprised for some reason at your capacity to reason this out. “You know what sucks about AI… is that almost everyone is using it… and it’s going to become hard for me to tell you all apart from it :(

@karpathy Yes, learning and understanding pixcels resembles/simulates human learning, vision based.

@karpathy I want help you with the image-input-only version of nanochat. Because it is how humans do it, we do not use tokens we use light, and the computer pixels!

@karpathy So the future is not text-to-text… but image-to-thought? 🤔

@karpathy how do you solve capital i vs small l through pure vision?

@karpathy Your PhD thesis was the first really good example I can think of as an example of a neural net based image-to-text pipeline (i.e., captioning natural images) for something other than OCR/handwriting recognition. I guess this isn’t all that different if you think about it.

@karpathy >side quest an image-input-only version of nanochat This is not the side quest. this is the main question you've talked about for years and enough public research exists to pull it off well. this would be a real improvement to the state of ai. a simple clean all-purpose nanoai

@karpathy Exactly but limited still. Not only text, then images, then video, then real world, but use new Conjunctures and previously unthought conceptions of the “real world” to train the new models

@karpathy OCR models are like translators for pixels, quietly bridging sight and language. What’s fun is how every small accuracy boost feels like teaching machines to read a little more like us. Maybe computer vision isn’t just seeing, it’s understanding.

@karpathy hi Andrej,dots means dots.llm1?

@karpathy Yep! This is a pretty famous paper in the tokenization and multi-lingual communities. Replacing BPE with rendering. https://arxiv.org/abs/2104.082...

@karpathy So the LLM is getting much more effective with the help of "photographic memory"...

@karpathy 💯There is a lot of garbage a text based LLM deals with when you have a tokenizer. Humans certainly don’t tokenize like an LLM does. So indeed why not process text the same way that a human does, by recognizing the shape of the letters and all the typography that comes with it.

@karpathy Should we start expending our “vision” models beyond the visible part of the spectrum? Isn’t this an important step if we are trying to have a more accurate understanding/representation of the world around us?

Share this thread

Read on Twitter

View original thread

Navigate thread

1/34