LiveCodeBench Pro sets a new standard for coding evaluation and is accepted to @NeurIPSConf 🧵 LiveCodeBench Pro evaluates end-to-end algorithmic reasoning under strict judges, real resource limits, and adversarial hidden tests so scores truly communicate coding capabilities.
2/ The challenge: LLMs are not elite at coding yet Competitive programming demands exact syntax and long-horizon reasoning: choosing the right data structures, proving complexity under tight time and memory budgets, and surviving adversarial edge generators. Many prior
3/ What LCB-Pro actually measures (end-to-end) We start when a contest round ends. The pipeline crawls the official statement, input generator, and judge, then freezes the original time and memory limits. We add a Codeforces-style hack phase plus in-house fuzzing to harden the
4/ Real contest problems, frozen limits & human-calibrated tiers Codeforces supplies a continuous stream of fresh problems across the full spectrum of difficulty, giving LCB-Pro a realistic distribution of topics and variants right after contests end. ICPC brings classic
5/ Submit and compare (apples-to-apples) Clone the repository, set up Python 3.12 and Docker, and implement a small adapter that invokes your model. Then run python http://benchmark.py to execute the full evaluation locally. The public leaderboard uses the exact same
6/ From score, to diagnosis, to improvement If you're training reasoning or code-gen systems, you need a benchmark that can’t be overfit with prompt tricks. LCB-Pro provides a clear and untainted image of coding gaps: long-horizon logic, brittle search, bad pruning, or poor
All the links 👉 Check out the live leaderboard standings and docs: http://livecodebenchpro.com 👉 Read the paper: https://www.alphaxiv.org/abs/2... 👉 Try LCB-Pro out (Toolkit & GitHub): https://livecodebenchpro.com/r...
@SentientAGI @NeurIPSConf Sentient keeps shipping
@SentientAGI @NeurIPSConf Congrats on the NeurIPS acceptance! It’s great to see real coding skills tested with strict rules and hidden challenges. This sets a strong benchmark for fair evaluation.
@SentientAGI @NeurIPSConf You create unique things that surprise us more and more!😎🔥
@SentientAGI @NeurIPSConf Interesting…
@SentientAGI @NeurIPSConf Keep building team
@SentientAGI @NeurIPSConf gSentient
@SentientAGI @NeurIPSConf Impressive work, can't wait to try LCB-Pro! 🔥🚀
@SentientAGI @NeurIPSConf boutta read it, give me a sec
@SentientAGI @NeurIPSConf Another top development from the team https://x.com/Ji1n11/status/19...
@SentientAGI @NeurIPSConf That’s interesting. Need to try it out in a bit
@SentientAGI @NeurIPSConf okay now drop more code
@SentientAGI @NeurIPSConf senti will rock the world!
@SentientAGI @NeurIPSConf Now I see why Sentient keeps winning awards
@SentientAGI @NeurIPSConf 太让人兴奋了
@SentientAGI @NeurIPSConf It turns out that our model is being tested specifically on LiveCodeBench Pro.:?
@SentientAGI @NeurIPSConf LiveCodeBench Pro sets a new standard for coding evaluation
@SentientAGI @NeurIPSConf Sentient keeps on shipping product after product!
@SentientAGI @NeurIPSConf 🦾
@SentientAGI @NeurIPSConf Sentient is built to sets all new standards for Open Source e AGI 🔥
@SentientAGI @NeurIPSConf LiveCodeBench Pro provides a precise and reliable evaluation of LLM coding performance, focusing on real contest problems and rigorous testing rather than simplified benchmarks. The approach ensures results reflect true algorithmic reasoning and practical coding ability.
@SentientAGI @NeurIPSConf Day by day big news and updates
@SentientAGI @NeurIPSConf Impressive! LiveCodeBench Pro is setting a new benchmark for coding evaluations 🧵
@SentientAGI @NeurIPSConf it’s a very interesting thanks for sharing
@SentientAGI @NeurIPSConf I think this one will ship in an almost perfect open source AI Model capable of evaluating coding under real conditions, and hidden tests make the results actually mean something. I'm curious to see how developers perform on this benchmark.😀😀😀
@SentientAGI @NeurIPSConf If LLMs get good here, that’s meaningful progress toward actual competitive-programming competence
@SentientAGI @NeurIPSConf Honestly impressed by how LCB grounds evaluation in real engineering challenges transparent, demanding, and truly reflective of problem-solving depth. Feels like a real leap forward for coding benchmarks.
@SentientAGI @NeurIPSConf Another big win for open-source AI! Sentient is literally setting the standards
@SentientAGI @NeurIPSConf This is the kind of benchmark the field’s been missing something that actually hurts a bit to fail. Real limits, real adversarial tests, no prompt magic. Finally, a mirror that shows what coding intelligence really looks like. gSenti team 🫶
@SentientAGI @NeurIPSConf You are an excellent team. Even if we say that two of the four different topics are ongoing, securing a place with these headlines in such an important place is a huge success. I sincerely congratulate you.


