Published: October 23, 2025
34
188
1.0k

I am super excited to share a new AI tool, Refine. Refine thoroughly studies research papers like a referee and finds issues with correctness, clarity, and consistency. In my own papers, it regularly catches problems that my coauthors and I missed. 1/

Image in tweet by Ben Golub

We’ve been beta testing it with frontier research in areas from social science to applied mathematics to computational biology to philosophy. Refine is already in use by editors and referees at top journals. 2/

Researchers find the comments thorough and helpful, exceeding the performance of top models like ChatGPT Pro. 3/

How does Refine work? First, it combines the strengths of several leading AI models. Second, it devotes many hours of compute to examining every detail of a paper. This makes for reports that save hours or days of researcher time. 4/

Another difference from many widely used AI products is that neither we nor our providers train models on the papers researchers upload. 5/

Here's some user feedback that has made us especially excited about sharing this tool with the wider world, and then a link to try it yourself. 6/

Image in tweet by Ben Golub

If you create a new account at http://refine.ink, you get a free preview that gives you a representative sample of comments on your own research. Please reply/QT/DM to share your feedback! 7/7

Image in tweet by Ben Golub

@ben_golub Just tested! It's great, but to be honest, a bit pricey!

@analisereal Would love to make it cheaper - exploring ways to do that. We have institutional subscriptions, which certainly make it cheaper for individual researchers ! 😅

@ben_golub congrats!

@IvanWerning thanks, Ivan!

@ben_golub Can it bin things into big glaring issues vs smaller ones? The marginal costs to fix everything keep going up.

@Bielsabub Greet feature request. Thanks!

@ben_golub Is this accessible somewhere?

@ben_golub How do you suggest people report using this when submitting their work to journals

@MaxJordan_N I don't think there are norms yet. An acknowledgment in the same section that you would acknowledge human feedback is certainly sufficient by any standard, and several authors have done it, but I don't think of it as necessary by any means until clearer norms emerge.

@ben_golub In your opinion, how does its feedback compare to the kind of reviews one normally gets from academic journals? My guess is better, although maybe also lacking something... which would make machine-augmented peer-review the way to go?

@joefrancis505 yes, that's exactly my take much more thorough in certain precision-heavy/boring details, which often highlight flaws that are wider-reaching, but generally much worse than top experts (for now) in assessing certain high-level issues about contribution and significance

@ben_golub Does it require full compilable latex or could you eg paste in one section? I’m curious to try it but I don’t have anything that’s like a full 10 page paper or something that length.

@npparikh It works with any PDF or other standard format It does work best for complete papers that are longer than the few pages that are very well handled by standard chatbots. But basically any academic text with significant reasoning extending across pages is fair game For

@ben_golub I tried it, on the free trial option. It is beautiful! No GPT-slop, accurate and thoughtful. A bit pricey perhaps, but a real option for that last step before one presses the button...

@joachim_voth Thank you!!

@ben_golub how could you possibly justify $50 for one review?

@andromedutch Try a review and let me know if you still have the question!

@ben_golub Congratulations, Ben!

@PeterBlairHenry Thank you, Peter!

@ben_golub Hi Ben. Great! where can I try it?

@ben_golub I see AI tools like this and ask, "How can productivity not be heavily accelerated by AI?" How long would it take you to find those issues or deal with faulty conclusions before you got back on track?

@ben_golub Refine could revolutionize how AI enhances academic rigor by automating referee like scrutiny. Spotting overlooked flaws in papers might just accelerate breakthroughs in AI research itself. 🚀

@ben_golub I tried it on one of my previously published papers ("Measuring LEDs"). Comments were interesting, mostly relating to ways to improve clarity. I would rate its performance as "on par" with human reviewers for everything except speed and bias. Speed was weeks faster. For bias,

@ben_golub I can second that this is exceptionally useful and is better than Gemini/ChatGPT in terms of catching issues!

Share this thread

Read on Twitter

View original thread

Navigate thread

1/33