Published: September 9, 2025
5
87
341

Latest genomic AI models report near-perfect prediction of pathogenic variants (e.g. AUROC>0.97 for Evo2). We ran extensive independent evals and found these figures are true, but very misleading. A breakdown of our new preprint: 🧵

Image in tweet by Nadav Brandes

We created a benchmark of ~250,000 pathogenic & benign variants. Unlike previous benchmarks, we evaluated performance by variant type. We broke down broad categories like ‘noncoding variants’ into specific annotations like intron, 3′ UTR and RNA gene.

Image in tweet by Nadav Brandes

Measuring model performance by variant type reveals a big anomaly. Evo2, for example, scores AUROC=0.975 on noncoding variants, but much lower on all specific types (e.g. 0.697 for splice, 0.903 for intron, 0.767 for 5′ UTR). Other models show similar pattern. What's going on?

It’s basically Simpson's paradox. To illustrate what’s happening, let’s look at Evo2 for splice & 5’UTR variants. Neither group shows good separation between pathogenic & benign variants, but splice variants get more damaging predictions & are much more likely to be pathogenic.

Image in tweet by Nadav Brandes

So when the two groups are merged, almost all pathogenic variants are splice and almost all benign variants are 5’UTR. You end up with almost perfect separation, just because the model knows to assign more damaging predictions to splice variants.

Image in tweet by Nadav Brandes

To show how serious this is, we included a simple rule-based baseline that only uses variant type information (no sequences, no AI). It achieves AUROC=0.944 across noncoding variants. The reported numbers suddenly look much less impressive.

Once you control for variant type, a clearer picture emerges. Reliable performance (AUROC>0.9) is achieved for missense, synonymous, non-splice intron, 3′ UTR & RNA gene variants. By contrast, stop-gain, start-loss, stop-loss, splice & 5′ UTR variants remain difficult.

Model comparison: -- GPN-MSA = most robust DNA model -- AlphaMissense = most robust protein model -- AlphaGenome & Evo2 = strong in some variant types, very unstable in others No single model is best across the board.

Lesson: near-perfect prediction of pathogenic variants across ‘all variants’ is an illusion. Variant-type-specific evals are needed to know when models are actually good and useful.

To ensure that future progress is meaningful, we provide the full benchmark & code. Check out our preprint: https://www.biorxiv.org/conten...

This work was done by three incredible students! @PoYu_Lin_NCKUH @BaiyuLu66681 @XueshenLiu

@BrandesNadav Nice work! Makes sense that you need completely different models for missense vs non-coding, splicing vs 3'UTR. We need much more experimental data! 250k variants is 0.01% of the genome

@BrandesNadav It’s fascinating to hear about the advancements in genomic AI models with such high AUROC values. While these metrics are impressive, your emphasis on their potentially misleading nature is crucial. Metrics need interpretation within the broader biological and clinical contexts.

@BrandesNadav Nice work, stratifying predictions by category is very helpful. By using ClinVar, the eval. conflates many levels of structural/functional impact (regulatory, transcript, protein, tissue, organ, organism). A variant may knock out a gene but benefit the organism (e.g. PCSK9) 1/2

Large AI models are reported to achieve high accuracy (AUROC) predicting pathogenic variants across the genome. A preprint reports that the predictions are based on splice variants. Using only this info (no sequences, no AI) achieves AUROC=0.944 across noncoding variants. 1/2

Image in tweet by Nadav Brandes

I'm really glad this re-evaluation was done but I just want folks to understand that when they see auROCs > 0.9 in any biological task, it is very rarely means they model is spectacular. It usually means the benchmark is severely flawed in some way. 1/

Ever wish you could hit "undo" on disease? 🩺🔄 https://www.nature.com/article... Most drug discovery asks: what does this perturbation do to cells? But we can also ask the reverse: which perturbations undo a disease signature and move cells back toward health? That's the idea behind

Image in tweet by Nadav Brandes

Vaccines are not magic. None are perfect. Some work better than others. Some have serious and occasionally lethal side effects. We don’t understand the downsides of vaccines perfectly. Not everyone needs every vaccine. Some are pushed for profit. The regulatory system around them

Share this thread

Read on Twitter

View original thread

Navigate thread

1/18