← Blog

Best AI Tools for Research: One Per Stage

September 18, 2026

AI Agents & ToolsProductivity

Most "best AI research tools" lists rank fifteen products against each other as if they were substitutes. They aren't. Research has five distinct stages, most tools cover one or two well, and the useful question is which tool for which stage.

The stages, and the strongest pick in each:

StageWhat you're doingStrongest pickFree?
DiscoveryFinding what exists on your topicSemantic ScholarYes
MappingSeeing how papers connectResearchRabbit, Connected PapersYes
ScreeningExtracting the same fields across many papersElicitPaid above an entry allowance
SynthesisReading deeply, asking questions of sourcesNotebookLMYes
CitationsManaging references and generating bibliographiesZoteroYes, open source

A reasonable full stack costs nothing: Semantic Scholar to find, ResearchRabbit to map, NotebookLM to synthesize, Zotero to cite. Paid tools earn their place at volume — systematic reviews across hundreds of papers — not at the start.

The rest covers what each stage actually needs, and the failure mode that matters more than any tool choice.

Discovery: Semantic Scholar

Semantic Scholar indexes well over 200 million papers with AI-generated summaries and a navigable citation graph, and it's free. For finding what exists on a topic, it's the strongest starting point available at any price.

Google Scholar remains the broadest index and is worth running in parallel — it catches grey literature, theses, and older work that Semantic Scholar misses. It has no AI layer to speak of, which is fine; discovery is a search problem more than a generation problem.

For preprints in physics, maths, computer science, and increasingly biology, arXiv is where the work appears first. Papers there haven't been peer reviewed, and treating a preprint as settled is one of the more common ways to get something wrong.

Mapping: ResearchRabbit and Connected Papers

The stage most people skip, and the one that changes how a literature review feels.

Both ResearchRabbit and Connected Papers take a paper you already have and build a visual graph of what it cites, what cites it, and what sits nearby in the citation network. You find the work that matters without knowing the right search terms in advance — which is exactly the problem when you're new to a field.

The specific value is catching the paper everyone in the field knows and nobody in your search results mentioned, because it uses different vocabulary than you did.

Screening: Elicit

When you have eighty papers and need the same five facts from each — sample size, method, effect size, population, limitations — that's extraction, and doing it by hand is where weeks disappear.

Elicit is built for exactly this. You define the columns you want, point it at a set of papers, and it fills the table with a link back to where in each paper it found the claim. That link back is the part that makes it trustworthy enough to use — you can check any cell against the source.

This is the one stage where paying is often correct. If you're doing a systematic review at real scale, the time saved is not marginal. For a literature review of a dozen papers, reading them yourself is faster than setting up the extraction.

Consensus sits nearby with a narrower job: you ask a yes-or-no research question and it shows how the weight of published evidence falls. Good for a quick orientation on a contested question, not a substitute for reading.

Synthesis: NotebookLM

NotebookLM is the best free tool for the "I have twenty PDFs and need to actually understand them" stage. You upload sources and ask questions grounded in those sources rather than in the model's training data, with citations back to the specific passage.

The grounding is the point. A general-purpose assistant answering from training data will produce a confident answer about your topic that may not reflect the documents in front of you. A tool constrained to your uploaded sources will tell you when the answer isn't there — which is the behavior you want.

General assistants like Claude and ChatGPT do belong in this stage too, for a different job: arguing with your interpretation, stress-testing an argument, rewriting a dense paragraph into something a non-specialist can read. Our comparison of Claude and ChatGPT covers which suits which of those.

Citations: Zotero

Zotero is free, open source, stores references, attaches PDFs, syncs across machines, and generates bibliographies in any style through a word processor plugin.

It has no AI features worth the name, and that's not a problem. Citation management is a solved database problem, and the solved solution is free. Spend the tool budget elsewhere.

The failure mode that matters more than the tools

Every tool above can produce a citation that doesn't support what it's attached to — or, in the worst case, doesn't exist. Fabricated references with plausible authors, real-sounding journals, and invented DOIs have appeared in submitted papers, filed legal briefs, and published articles.

The tools that retrieve first and generate second — Semantic Scholar, Elicit, NotebookLM, Consensus — are much safer here than a general chat assistant, because they're constrained to real documents. Safer is not immune. A retrieval tool can still attach a real citation to a claim the paper doesn't make.

The only reliable defense is unglamorous: open the source and check that it says what the tool says it says. Not every citation in a long review, but every citation you rely on. Our guide to AI source finders goes deeper on which tools genuinely retrieve and how to run that check quickly.

Treat this as the cost of using the tools, not as an argument against them. The time saved on discovery and extraction is real, and it comfortably funds the verification.

Building a stack that persists

The awkward part of a five-tool stack is that nothing carries between the tools. Your screening criteria live in Elicit, your reading notes in NotebookLM, your references in Zotero, and the process you worked out for your last review lives in your head — so the next project starts from scratch.

That's the gap Taku works on: instead of re-running a process from memory, you mirror an AI app or workflow that already does it, run it in your own desktop workspace, and keep it for the next project. If your research process is a set of steps you repeat and re-explain every time, the free app library is the place to start. Taku is in Beta, and the Mac app is available now.

Key points

  • Research has five stages and no single tool covers them all. Match the tool to the stage.
  • A capable free stack exists: Semantic Scholar, ResearchRabbit, NotebookLM, Zotero.
  • Paid extraction tools like Elicit earn their cost at volume, not on a small literature review.
  • Mapping tools find the important paper your search terms missed. Don't skip that stage.
  • Retrieval-based tools are far safer than general chat assistants for anything with a citation attached.
  • Verify every citation you rely on against the source. That check is the price of the time saved.

FAQ

What's the best free AI tool for research?

For discovery, Semantic Scholar. For synthesis, NotebookLM. For citations, Zotero. All three are free and together they cover most of a literature review.

Can AI write my literature review?

It can draft sections and it will produce claims that need checking against sources. The synthesis — the argument about what the literature collectively shows — is the part that carries your judgment, and it's also the part that's hardest to verify if you didn't do it yourself.

Do AI research tools make up citations?

General chat assistants can. Tools built on retrieval, like Elicit, NotebookLM, and Semantic Scholar, are constrained to real documents and are much safer — though a real citation can still be attached to a claim the paper doesn't make. Check what you rely on.

Is Elicit worth paying for?

At volume, usually yes — a systematic review with hundreds of papers to screen is exactly what it's for. For a dozen papers, reading them yourself is faster than setting up the extraction.

What's the difference between Semantic Scholar and Google Scholar?

Google Scholar has the broader index and catches grey literature. Semantic Scholar has AI summaries and a much better citation graph for exploring connections. Run both.

Can I use ChatGPT or Claude for research?

For thinking, drafting, and rewriting, yes, and they're good at it. For finding sources, use a tool that retrieves from a real index — a general assistant answering about literature from training data is where fabricated citations come from.

What about tools for non-academic research?

The same shape applies: something that searches a real index, something that reads your own documents, something that keeps your notes. The sources change; the stages don't.