AI models are dramatically improving at IQ tests (70 IQ → 120), yet they don't feel vastly smarter than two years ago. At their current level of intelligence, rehashing existing human writings will work better than leaning on their own intelligence to produce novel analysis.
This is also an explanation for why AIs can't come up with good jokes yet.
@DanHendrycks "they don't feel vastly smarter than two years ago" is actually braindead have you used anything past gpt4o? if you have and still stand by this, i have no words
@DanHendrycks > yet they don't feel vastly smarter than two years ago. This is the stupidest thing I've read in a while.
@DanHendrycks it's unclear to me that IQ tests mean the same thing for an LLM as they do for a human
@DanHendrycks completely disagree with this. they feel vastly more capable. much more accurate in analysis and capable of problem solving, providing first draft analysis and beyond. run a gpt4 class model locally and compare it to 3.7T or 2.5P. worlds apart.
@DanHendrycks You don't think they feel vastly smarter than two years ago? They definitely feel that way to me.
@DanHendrycks llms are vastly smarter than two years ago, how about you go back to gpt 3.5 being the only model.
@DanHendrycks They feel vastly smarter than they did two years ago. I can do things with them now that I couldn’t possibly have done two years ago.
@DanHendrycks My view about all this is extremely different; would you be willing to do a podcast about it, since it requires some depth and is important?
@DanHendrycks What are you talking about? Legacy ChatGPT-4 from two years ago is a useless, dumbass in comparison to o3 and o4-Mini. 🤷🏽♀️
@DanHendrycks This seems very narrow and reductive. Humans can produce non-slop without needing to cross some mythical IQ threshold. Or jokes, for that matter. If the AI is unable to do so at sub-genius levels, clearly it's missing something other than IQ.
@DanHendrycks The minimum intelligence to contribute good ideas to a lot of fields is low for most stuff I don't think this is the barrier. It's true that a power law apply to the publication rates/well cited ideas but most fields are full of midwits who produce good idea and contribute
@DanHendrycks But nobody is even trying to sample "novel analysis" properly. People produce novel results after they spend years on their own unique trajectory. Also we have billions of people sampling all the time. And, frankly, o3 does feel like it's 150 IQ tbh.
@DanHendrycks Trust me, they are indeed smarter. If you can't see it, then you've got to question your prompting skills.
@DanHendrycks Have we collectively forgotten about how bad models were 2 years ago compared to now? No need for gaslighting.
@DanHendrycks This is an interesting take. To try and rephrase what I think you said: AI has gone from 70->120IQ, & anything in that range succeeds by quoting the work of others more than originality. There will be a threshold soon where original AI thought is more compelling. Thanks!!
@DanHendrycks Why is o4 high iq so low 😭
@DanHendrycks >“yet they don't feel vastly smarter than two years ago” Couldn’t disagree more. It’s not even close to where we were 2 years ago.
@DanHendrycks we have models already cruising past 130+ IQ and my model just helped backfill for loss to launch PHP+ to WASM transformer using Node. the web is about to get a LOT more optimized, what we need are more namespaces with CA for trusted idea perfection if you will, this will lead to
@DanHendrycks AI improvement curve is steep threshold nearing.
@DanHendrycks cool
@DanHendrycks @threadreaderapp @Twtextapp @unrollthread unroll @threader_app compile ========================================
@DanHendrycks Al is here to improve lots of things
@DanHendrycks > they don't feel vastly smarter than two years ago. How can people make such confident claims? Just chat with GPT-4 and then with o3, 4.5, or even 4o. The difference in quality and intelligence is obvious.
@DanHendrycks That's because IQ tests are made to measure intelligence in humans, not in machines (whether they even do the former in a good way is a separate question). A normal calculator does math faster than like 99,9% of humans, but is not especially smart despite that.
@DanHendrycks It's been said that a lower-IQ person can't perceive a difference between the intelligences of two different higher-IQ people. Not wanting to brag but I've noticed a big jump in AI model intelligence.
@DanHendrycks There is still the GIGO (Garbage In - Garbage Out) problem. A major portion of what current AIs believe is lies, opinion, & disinformation. Just like human intelligence, AI has difficulty overcoming cognitive dissonance, normally choosing to ignore it, just like most humans.
@DanHendrycks You can't "feel" the IQ of most people. And most people of very high IQs produce slop and just rehash what came before. Intelligence has no intrinsic link to originality or insight.
@DanHendrycks Wtf are you smoking?? They don’t feel vastly smarter than they did two years ago???? Now, I’m not saying we have super intelligence, but damn.
@DanHendrycks Nothing to do with your post, I just thought this was funny
@DanHendrycks The uncomfortable corollary is that the vast majority of the human race produces slop.
@DanHendrycks I don’t think you have a very strong understanding of AI models. AI models from 2 years ago struggle with the simplest of complex tasks, Current models feel intelligent and useful. If you actually cannot tell the difference between 70 and 120iq I feel really sorry for you
@DanHendrycks AI feels way smarter on paper — the IQ tests, the benchmarks, all that. But when you use it, it still kinda spits out the same rehashed stuff. Real insight doesn’t scale with IQ in a straight line. It jumps at a certain point. And we’re probably not there yet. Maybe 10 more IQ
@DanHendrycks Expect that to come via routes similar to the Absolute Zero Reasoner (AZR; from yesterday). Their own Q, their own A, three-way learning of the (x,y,f) triplet making up y=f(x): assuming any of the two in the triplet, learning the third, in every three-ways combo




