Published: October 22, 2025
31
199
1.3k

After an interview with @karpathy, everyone is talking about what AI agents can/can't do. But an opinion without data is just a hypothesis. So, I tested 3x185 workflow executions for a market researcher agent. The results have shocked me🧵

Image in tweet by Paweł Huryn

I tested three variants: I. LLM Workflow: No agency, the entire logic carefully orchestrated. What was expected: - An LLM workflow was 2x faster (the same model) compared to an AI Agent. - An LLM workflow consumed 12x less tokens to an AI Agent. 3/185 "errors" are minor

Image in tweet by Paweł Huryn

II. Agentic Workflow: Deterministic logic moved to the orchestration layer. More time, more tokens. 100% task success. GPT-5 (a reasoning model) consumed less tokens than GPT-4o due to better compression. None of this was surprising. But then:

Image in tweet by Paweł Huryn

III. AI Agent: Full autonomy without steps to take, just an objective I were staring at the screen. An AI agent without predefined reasoning steps succeeded 185/185 times (100%).

Image in tweet by Paweł Huryn

This is different from my previous observations for the same models: https://x.com/PawelHuryn/statu...

Conclusions & learnings: 1. For simple use cases, we can already achieve 99%+ reliability 2. A verifier agent with a high TPR would push it even further 3. For complex or critical processes, you still need orchestration 4. Orchestration is faster, cheaper, and more reliable

@karpathy might be right. We might need 10 years to achieve true AI intelligence. But autonomy and reliability for most processes seem more like ~12 months away. Agree? Disagree? Let me know in the comments. P.S....

A. Free n8n templates I used for testing: https://www.productcompass.pm/...

B. Enjoy this? - Follow me @PawelHuryn for deep researched AI & PM - Share this thread with others I appreciate it! https://x.com/PawelHuryn/statu...

@PawelHuryn @karpathy I’m writing a script that will block anyone who starts a thread with the words “shocked” “blew my mind” “freaking me out” “you won’t believe” or any derivative thereof.

@wayne_culbreth @karpathy I hear you, but that way you might eventually block everyone. People are shocked sometimes. Nobody should be "shocked" or see "crazy" things in more than a few % of their post.

@PawelHuryn @karpathy Love when people actually work with data instead of broad generalizations

@PawelHuryn @karpathy But what about use cases that you can’t do with a that kind of calls, What about scenarios when you need to read lot of data and select only the relevant and give it to the master agent? If you can solve a problem without autonomy you always should! Any how, I enjoyed the

@PawelHuryn @karpathy The ecosystem is still young. We have a lot of problems to solve before they reach their potential.

@PawelHuryn @karpathy Nothing spices up AI debates like someone actually running the numbers. Opinions crash quick when data enters the chat. Maybe the real intelligence isn’t artificial, it’s statistical.

@PawelHuryn @karpathy Pessimists sound smart. Optimists make money.

@PawelHuryn @karpathy Interesting 🙌🏻

@PawelHuryn @karpathy Thanks for providing this great insight.

@PawelHuryn @karpathy @grok explain this thread to me in simple words so that the concepts are clear.

@PawelHuryn @karpathy This is exactly the kind of data-driven approach we need! Testing 3x185 executions gives real statistical significance. The jump from 94% to 100% success while reducing tokens from 18K to 4.8K shows agents are both more effective AND efficient. Great work!

@PawelHuryn What tool are you using to run workflows?

@PawelHuryn @karpathy Interesting. However, I thought that Agentic AI models are more autonomous than regular AI Agents. It’s supposed that Agentic AI is the next evolution.

@PawelHuryn @karpathy This comparison is interesting Pawel, I use Zapier agents and set them up with a goal for my marketing. Where a year ago the results would sometimes be off in terms of sticking to the context and goals. I now see higher quality outcomes and more self solving. #AIAgent

@PawelHuryn @karpathy Finally, real benchmarks instead of speculation. Agent discourse needs more data-driven validation like this to separate hype from actual capability.

@PawelHuryn @karpathy @PawelHuryn Thank you so much, this is really valuable work 🙌. I confirm, I came to similar conclusions a few days ago using approaches B and C for a PoC I'm working on. I love people who use a scientific approach to validate hypotheses.

@PawelHuryn @karpathy Which workflow took longer to put together?

@PawelHuryn @karpathy GPT-5 Token Usage ratios {1, 3.5, 6.5} (actual tokens {4.5K, 15.5K, 30K}) I. LLM Workflow: No agency,entire logic carefully orchestrated II. Agentic Workflow: Deterministic logic moved to orchestration layer III. AI Agent:Full autonomy without steps to take, just objective

Image in tweet by Paweł Huryn

@PawelHuryn @karpathy GPT-5 Token Usage ratios {1, 3.5, 6.5} for the three types of workflows {LLM Workflow, Agentic Workflow, AI Agent} most ( I would say 75%+) enterprise workflows are pre-defined/pre-determined, the first two types will do the job cheaper (less Tokens) https://x.com/PostPCEra/status...

@PawelHuryn @karpathy What exactly did your data show about the agent's performance after those 555 workflow tests?

@PawelHuryn @karpathy Hey Pawel, i‘m working on similar research agent for a lage scale automotive company for my thesis, can i DM you?

Share this thread

Read on Twitter

View original thread

Navigate thread

1/32