Published: October 21, 2025
78
343
2.2k

I used ChatGPT to solve an open problem in convex optimization. *Part I* (1/N)

Problem statement: (2/N)

Image in tweet by Ernest Ryu

The proof, cleaned up and typed up by me: (3/N)

Image in tweet by Ernest Ryu
Image in tweet by Ernest Ryu

Key steps of the proof produced by ChatGPT: https://chatgpt.com/share/68f8... (4/N)

My reaction: ChatGPT was really effective at accelerating my progress. This work took about 12 hours, spread over 3 days. In hindsight, the proof is really simple. (5/N)

But I iterated through so many other strategies that didn't pan out, and ChatGPT crucially helped to quickly explore and eliminate those dead-end approaches. Also, the key successful steps were suggested by ChatGPT. (6/N)

ChatGPT did not produce the proof in a single prompt. The process was highly interactive. It generated many arguments, roughly 80% of which were incorrect. (7/N)

Yet some were genuinely novel to me. Whenever I recognized a novel idea, whether correct or only partially so, I distilled the key insight and prompted ChatGPT to develop it further. (8/N)

My contribution: - Filtering out incorrect arguments and accumulating a set of correct facts. - Identifying promising new lines of reasoning and guiding ChatGPT to explore them further - Recognizing when a strategy had been fully explored and deciding when to move on. (9/N)

ChatGPT's contribution: - Producing the final proof argument. - Significantly accelerating my (or our) exploration of the many dead-end arguments, rapidly ruling out approaches that did not work. (10/N)

Plans forward: In my view, this result is already publishable in a respectable optimization theory journal. However, I would like to flesh it out further. Here are the next steps I plan to take. (11/N)

*Part II (forthcoming)* I'll attempt to generalize to the ODE with r>0. The convergence and divergence behavior is already known for r≤1 (divergence) and r>3 (convergence). I'll aim to extend our analysis to the intermediate regime, r∈(1,3). (12/N)

Image in tweet by Ernest Ryu

*Part III (forthcoming)* We'll see if I can translate the argument to prove convergence of the discrete-time counterpart, namely of Nesterov's accelerated gradient method. (This is actually the main open problem.) (13/N)

Image in tweet by Ernest Ryu

Roadblock: However, I've run out of ChatGPT Pro queries. I'm on the expensive pro-tier plan, but I've already used up my allowance, and it doesn't refresh until next month. Could someone from OpenAI help me out with this? (14/N) @SebastienBubeck @kevinweil

@SebastienBubeck @kevinweil Conclusion: I will continue to develop this project and publish the results in a respected optimization theory journal. I'll share updates and future parts (II and III) here on Twitter as the work progresses. (15/N)

@SebastienBubeck @kevinweil ChatGPT is now at the level of solving some math research questions, but you do need an expert guiding it. This exercise was a lot of fun and was highly productive. I also feel I'm getting better at prompting ChatGPT. I'll also try other open and unsolved problems. (16/N, N=16)

@ErnestRyu Most people at @OpenAI can and will help you with a pro subscription. Please DM me if you haven't been helped. Congrats. Nice result.

@hyhieu226 @OpenAI Thanks! Seb and Kevin reached out, so I should be able to proceed with GPT5 Pro credits.

@ErnestRyu I inputted these pics into @HarmonicMath Aristotle and it produced a formally verified proof of the claim in Lean 4: https://gist.github.com/llllvv...

@ErnestRyu Could you be more precise on which ChatGPT ideas you found "novel"? It looks like a pretty standard proof.

@DocOctagonical ChatGPT generated many interesting leads that ended up being dead ends. I really appreciate those, but I'm not showing the failed attempts.

@ErnestRyu the flashbacks !

Image in tweet by Ernest Ryu

@ErnestRyu @abeirami Very cool! Thanks for sharing

@ErnestRyu LLMs are a tool, and you know how to use it! That’s exactly how it should be used! I suspect the 80% wrong content will go down rapidly with better versions, eg GPT-6, GPT-7!!! But yeah, as a means of cheap stochastic exploration of paths it’s exactly right!

@ErnestRyu "It generated many arguments, roughly 80% of which were incorrect." In other words, it threw shit at the wall while you guided it towards an answer.

@ErnestRyu ChatGPT Pro is a hot mess. It does OK for some things, but then it hallucinates others. The results are often very plausible looking, but once you dig deeper, you will discover they are false. Still, it’s better to have it than to not have it. 🤷‍♂️

@ErnestRyu This is cool, but I’m curious: how much interest did this problem have? I see there’s 2 references stating this problem, but is it “well-known” to experts? Just asking because I’m rather surprised by the length of the proof (usually “known” open problems in NT have >30 page pfs)

@ErnestRyu I combined Grok 4, Gemini 2.5 Pro to solve the Hilbert–Pólya conjecture. ChatGPT-5 later checked it. ✅ https://x.com/grok/status/1978...

@ErnestRyu From personal experience Gemini 2.5 Pro has also been good at back and forth for proofs, and took tons of queries to run out of budget. Ps don’t give any of them real analytic geometry problems - they are crap at it

@ErnestRyu Very cool. I wonder if LLMs could be trained to do this type of filtering itself that you did, to move closer to an AI system that can both generate novel ideas and then select the best ones for moving forward. I guess this might be possible right now with a loop?

@ErnestRyu I’ve had the same experience (though with the “expertise” of a senior PhD student). I tend to describe it as a directionally correct high entropy idea generator. It might also have to do with the way I use it and it might be fine tuned differently for others

@ErnestRyu This post basically demonstrates perfectly well why AI would make GDP sky rocket. You have a problem statement, AI quickly removes all dead ends. Human SME verifies, from 2 weeks to 2 hours. EE, materials science, math, risk models, legal proposals, everything.

@ErnestRyu It would be nice to formalize the proof as a lean lang!

Share this thread

Read on Twitter

View original thread

Navigate thread

1/35