Published: May 16, 2024
154
455
2.6k

In the late 1960s top airplane speeds were increasing dramatically. People assumed the trend would continue. Pan Am was pre-booking flights to the moon. But it turned out the trend was about to fall off a cliff. I think it's the same thing with AI scaling — it's going to run

Image in tweet by Arvind Narayanan

By 1971, about a hundred thousand people had signed up for flights to the moon https://en.wikipedia.org/wiki/...

You may have heard that every exponential is a sigmoid in disguise. I'd say every exponential is at best a sigmoid in disguise. In some cases tech progress suddenly flatlines. A famous example is CPU clock speeds. (Ofc clockspeed is mostly pointless but pick your metric.)

Image in tweet by Arvind Narayanan

There are 2 main barriers to continued scaling. One is data. It's possible that companies have already run out of high-quality data, and that that's why the flagship models from OpenAI, Anthropic, and Google all have strikingly similar performance (that hasn't improved in > 1y).

What about synthetic data? There seems to be a misconception here — I don't think developers are using it to increase training data volume. This paper has a great list of uses for synthetic data for training, and it's all about fixing specific gaps and making domain-specific

The 2nd, and IMO bigger barrier to scaling is that beyond a point, scale might lead to better models in the sense of perplexity (next word prediction) but might not lead to downstream improvements (new emergent capabilities).

This gets at one of the core debates about LLM capabilities — are they capable of extrapolation or do they only learn tasks represented in the training data? It's a glass half full / half empty situation and the truth is somewhere in between but I lean toward the latter view.

So if LLMs can't do much beyond what's seen in training, at some point it no longer helps if you have more data because all the tasks that are ever going to be represented in it are already represented. Every traditional ML model eventually plateaus; maybe LLMs are no different.

My hunch is that not only has scaling already basically run out, this is already recognized by teams building frontier models. If true, it would explain many otherwise perplexing things (I have no inside information): – No GPT-5 (remember: GPT-4 started training ~2y ago) – CEOs

In my AI Snake Oil book with @sayashk (https://www.aisnakeoil.com/p/a... we have a chapter on AGI. We conceptualize the history of AI as a punctuated equilibrium, which we call the ladder of generality (which doesn't imply linear progress). LLMs are already the 7th step in our ladder; an

OpenAI released GPT-3.5 and then GPT-4 just a couple of months later (even though the latter had been in development for a while). This historical accident had the unintended effect of giving people a greatly exaggerated sense of the pace of LLM improvements, and led to a

@random_walker If there was some shared mechanism of action then this analogy would be insightful, but opening with "IT'S A SIGMOID NOT AN EXPONENTIAL LOOK AT THIS OTHER SIGMOID TRUST ME IT MUST BE A SIGMOID" isn't contributing anything to the discussion

@random_walker shocked that the author of “AI Snake Oil” thinks this 😱

@random_walker Airplanes have to contend with physical laws. What's the parallel? Only realistic explanation so far with scaling laws could be that the amount of capital required for larger training runs becomes hard to coordinate and the gains are only ~log in spend.

@random_walker Interesting. However, for plane speed there were natural and understood constraint, like air friction, or, ultimately, the speed of light.

@random_walker Interesting. What do you think the chances are of a 1e27 FLOP training run (~100x the scale of GPT-4) by end of 2027?

@random_walker And average speed of commercial aviation went down in last 20 years (from @beenwrekt substack on control)

@random_walker Honestly I think this is true but not because of what the tweet says; the physics of flying faster are more difficult than those around scaling AI. Both, however, have and will be hamstrung by regulation — supersonic flight over the U.S. was effectively banned and so too will

@random_walker The first transcontinental jet flight (LA-NYC) by a Boeing 707 in 1959 (nearly 65 years ago) was actually quicker than current scheduled flights.

@random_walker I booked marked this... you are so wrong 🤣🤣

@random_walker I wouldn’t put my money against ilya, ever

@random_walker I expect emerging forms of federated learning (will we even call it FL? as this happens) will help let AI continue scaling both on the data and compute side.

@random_walker Have you heard of the Speed of Sound?

@random_walker Easy way to test OOD capability is to construct a new eval for the areas you care about and keep the true test set private, with unique questions you know aren’t in the training data bc you made them. My sense is that it’s more the glass half full, that’s there’s more inductive

@random_walker Maybe. Current architectures are very data inefficient, nowhere near theoretical bounds. There are starting to be serious experiments with non-auto regressive approaches. We’ll see what they bring.

@random_walker It's a great analogy! Regulators banned supersonic flight, so commercial incentives for R&D stopped. Progress topped There are currently ~700 bills in state governments impacting AI. They range from reasonable to Loony Tunes History repeats?

@random_walker "640 KB of memory ought to be enough for anybody." - Bill Gates (1981) “By 2005 or so, it will become clear that the Internet’s impact on the economy has been no greater than the fax machine’s.” - Paul Krugman (1998)

@random_walker Speed stopped increasing because fuel consumption increases with velocity cubed, it’s too expensive to fly faster in the atmosphere. The only thing limiting the size of AI models is the return on investment. Can the income stream to a given size AI model provide sufficient return

@random_walker Absolutely brilliant take, the ability to conjure something uniquely new from scratch is something that LLMs lack, as the data on they train is already represented. Though offcourse, they can fill the gaps and logics that are missing in the hay of all that data.

@random_walker Supersonic flight was banned in 1973? What a coincidence.

@random_walker The scaling stopped because the gov banned civilian supersonic overland flights, it has nothing to do with the tech.

Image in tweet by Arvind Narayanan

@random_walker This viewpoint offers a counterbalance to the prevalent hype and doomsday predictions surrounding AI, emphasizing the importance of considering the possibility of natural limits to growth and innovation in this field. It's a refreshing reminder to remain grounded in realistic

@random_walker Whatever you do, there’s a local maximum.

@random_walker It is likely LLMs will plateau due to absence of novel text data and scale can't solve that. But there is still potentially infinite video and audio data, so you can continue scaling up.

@random_walker Ai made a leap when the CNN architecture came about. Ai made a leap with Attention Is All You Need. I see no reason to believe those were the last creations.

Share this thread

Read on Twitter

View original thread

Navigate thread

1/37