Elicit: AI Research Assistant, Systematic Review Automation, and the Accuracy Gap Explained
Updated: Sep 9
Elicit is an AI research assistant built specifically for academic literature: it searches more than 138 million papers using semantic similarity rather than keyword or Boolean matching, extracts structured data into comparison tables rather than returning summaries alone, and automates the screening and extraction steps of a formal systematic review. Spun out of Ought, a nonprofit machine-learning research lab, in 2023 as a public benefit corporation, Elicit has raised $31 million total, including a $22 million Series A co-led by Spark Capital and Footwork.
The product is deliberately narrow — it explicitly cannot help with current events, market data, or any topic whose evidence base isn't academic literature, since its database is indexed periodically rather than searched live. For a researcher evaluating Elicit, the deciding factor is whether automated systematic-review screening and extraction performs reliably enough for actual methodological use, given that Elicit's own validation numbers and an independent academic study of the same workflow report meaningfully different accuracy results.
··········
HOW SEMANTIC SEARCH AND STRUCTURED EXTRACTION ACTUALLY WORK.
Concept-based paper retrieval, tabular data extraction, and a PRISMA-auditable review workflow define the mechanism beyond a search bar.
Elicit's search finds conceptually relevant papers even when they don't use a query's exact terminology, which is a structural difference from a traditional academic database that requires precise Boolean operators matching paper metadata — a paper about a mechanism described with different vocabulary than the search terms still surfaces under semantic search, where it would be invisible to keyword matching. Once papers are retrieved, Elicit extracts structured data into a table rather than leaving a researcher to read and manually tabulate each one — comparison columns for methodology, sample size, or findings populate automatically, scaling up to 40 extraction columns at the Enterprise tier.
The Systematic Review workflow formalizes that extraction into something closer to a defensible research process: it screens papers against inclusion and exclusion criteria, logs the reason for each exclusion, and produces a PRISMA-auditable trail with criterion scores and supporting quotes — the documentation a peer-reviewed systematic review actually needs to justify its paper selection, rather than an informal reading list. Pro-tier accounts can screen up to 5,000 papers through this workflow; Enterprise scales to 40,000. A newer Research Agent runs a broader automated research workflow across these capabilities rather than requiring a researcher to chain search, extraction, and screening manually.
........
Component | Mechanism | Function |
|---|---|---|
Semantic search | Concept-based retrieval across 138M+ papers | Finds relevant papers regardless of exact query terminology |
Structured extraction | Tabular data pulled from papers automatically | Scales to 40 columns (Enterprise), avoiding manual tabulation |
Systematic Review workflow | PRISMA-auditable screening with exclusion reasons | Produces a defensible paper-selection record for formal reviews |
Screening capacity | Up to 5,000 papers (Pro), 40,000 (Enterprise) | Matches review scale to a project's actual literature volume |
Research Agent | Automated multi-step research workflow | Chains search, extraction, and screening without manual handoffs |
Enterprise controls | SSO, SAML, 2FA, single-tenancy, no training on data | Meets institutional security and data-handling requirements |
........
··········
WHY THE GAP BETWEEN ELICIT'S OWN ACCURACY CLAIMS AND INDEPENDENT TESTING MATTERS FOR SERIOUS USE.
A 59-percentage-point difference between controlled benchmark conditions and real-world search strategies is the single most important number to understand before relying on Elicit for a formal review.
Elicit validated its screening pipeline against 994 Cochrane systematic reviews and reported 96.9% abstract-screening sensitivity, a figure the company says exceeds single-reviewer human performance, alongside 95.0% search recall and 99.5% full-text paper-level recall. Those numbers, on their own, would suggest the tool performs at or above the reliability bar formal evidence synthesis requires.
An independent 2025 study published in Cochrane Evidence Synthesis and Methods (Lau et al.) tested the same screening task under a different condition — using search strategies structured the way real systematic reviewers actually query the literature, rather than the setup Elicit's own benchmark used — and found sensitivity dropped to 37.9%. That's a 59-percentage-point gap between the vendor-reported figure and independent field-condition testing, and it illustrates a distinction that extends beyond this one tool: a benchmark run under controlled conditions measures the model's best-case capability, not necessarily its performance when a working researcher applies their own search strategy rather than the one the benchmark was built around. For casual literature discovery, that gap may not change much in practice. For a systematic review intended for peer-reviewed publication, where screening sensitivity is a methodological claim reviewers will scrutinize, treating either number as settled without checking current, independent validation would be a mistake in either direction.
··········
HOW ELICIT'S PRICING COMPARES TO OTHER AI RESEARCH TOOLS.
Elicit costs more than lighter research assistants, but its extraction depth and review workflow are also functionally distinct from either lighter alternative.
Elicit's free tier includes unlimited semantic search across its full paper database with a limited number of automated reports per month; Plus, at roughly $10–12 per month, adds exports and clinical trial database access; Pro, at roughly $42–49 per month, unlocks the Systematic Review workflow and Research Agent; Team pricing runs around $79 per seat per month with a two-seat minimum, and Enterprise is quoted directly.
........
Tool | Core capability | Starting price |
|---|---|---|
Elicit | Semantic search, structured extraction, systematic review workflow | Free; Plus ~$10–12/mo; Pro ~$42–49/mo; Team ~$79/seat/mo |
Consensus | Evidence summaries from academic papers | Pro $9.99/mo |
SciSpace | Paper reading assistance | Premium from $12/mo |
........
Neither Consensus nor SciSpace offers Elicit's extraction-table depth or the systematic-review screening workflow specifically, which is why the price gap reads smaller in practice than the sticker prices alone suggest — a researcher comparing Elicit's $42–49 Pro tier against Consensus's $9.99 is comparing a systematic-review production tool against an evidence-summary tool, not two versions of the same capability at different prices.
··········
THE DECISION RULE FOR EVALUATING ELICIT.
Elicit earns its price for researchers running structured literature work where semantic search and automated extraction save real time over manual reading and tabulation — for that use case, the Plus tier alone often justifies itself against the alternative of doing the same work by hand. The Pro tier's Systematic Review workflow is a different commitment: it's built for reviews with methodological rigor requirements, and anyone relying on it for a formally published review should independently verify current screening accuracy rather than citing either Elicit's own Cochrane benchmark or the Lau et al. field-condition figure as the final word, given how far apart those two numbers already sit and how much a review's credibility depends on that specific claim. Anyone whose research need includes current events, market conditions, or any non-academic evidence base should look elsewhere entirely — Elicit's periodically indexed academic database, by design, doesn't cover that territory regardless of which accuracy figure turns out to be closer to the truth.
·····
FOLLOW US FOR MORE.
·····
DATA STUDIOS
·····




