Ai2 launches AstaBrief 8B open model for scientific reports with 3.5x faster generation

Ai2 has released AstaBrief 8B, an open-weights language model designed specifically to transform research questions and retrieved scientific literature into structured reports with citations.
The model powers the new Fast mode in Asta, Ai2's platform for scientific research. Across the complete Asta report-generation pipeline, Fast mode averages 51.1 seconds per report, compared with 178.5 seconds for the existing Claude-powered Thinking mode, making the new pipeline approximately 3.5× faster.
AstaBrief is based on Qwen3-8B and was specialized through supervised fine-tuning and direct preference optimization rather than reinforcement learning. Ai2 is releasing not only the model weights but also training data and an example workflow for generating reports from researchers' own PDFs.
The architecture targets a specific problem in AI-assisted science: producing useful literature syntheses while preserving citation grounding, evidence traceability and the scope of the underlying research findings.
··········
ASTABRIEF 8B AT A GLANCE
........
Specification | AstaBrief 8B |
Parameters | 8 billion |
Base model | Qwen3-8B |
Primary task | Scientific report generation |
Inputs | Research question + retrieved literature excerpts |
Output | Structured report with citations |
Training | SFT + DPO |
Initial research-query pool | 90,000 queries |
SFT examples after filtering | 47,000 |
Asta operating mode | Fast |
Average Fast-mode report time | 51.1 seconds |
Thinking-mode comparison | 178.5 seconds |
Pipeline speed difference | ~3.5× faster |
Weights | Open |
Training data | Released |
License | Apache 2.0 |
........
AstaBrief is not designed as a general-purpose replacement for frontier language models. Its training and serving pipeline are optimized around a narrower task: receiving already retrieved scientific evidence and converting it into a cited synthesis.
··········
ASTABRIEF GENERATES THE ENTIRE SCIENTIFIC REPORT IN ONE PASS
AstaBrief changes the report-generation architecture used by Asta.
The Claude-powered Thinking pipeline performs several intermediate operations before producing the final document, including processing retrieved evidence, organizing material and generating sections through a multi-stage workflow.
AstaBrief is trained to receive the research query and relevant retrieved snippets directly and generate the complete report in a single pass.
Removing intermediate summarization, clustering and section-by-section generation substantially reduces the amount of inference required for each report.
Across the complete pipeline, Ai2 measured average generation times of 51.1 seconds for Fast mode and 178.5 seconds for Thinking mode.
Data Studios calculation: the reduction from 178.5 to 51.1 seconds eliminates approximately 71.4% of the report-generation time, while the Thinking pipeline requires about 3.49× as long as Fast mode under Ai2's measurements.
The comparison measures the complete Asta pipelines rather than raw token-generation speed between two isolated language models.
··········
AI2 TRAINED THE MODEL FROM REAL SCIENTIFIC RESEARCH QUERIES
AstaBrief's post-training dataset began with queries submitted by researchers through Ai2's scientific research systems.
Ai2 analyzed a much larger collection of usage data and filtered the available queries for quality, relevance and privacy. Beta-tester and automated traffic was removed, alongside queries that were too short, non-scientific, non-English or contained personal information.
The filtering process produced approximately 90,000 research-focused queries.
Ai2 then used its existing multi-stage scientific report pipeline to generate training targets. The systems involved in creating those reports included Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1.
After further quality filtering, Ai2 retained approximately 47,000 examples for supervised fine-tuning.
The resulting model therefore learns the report-generation task from examples derived from realistic scientific questions rather than exclusively from synthetic benchmark prompts.
··········
CITATION DENSITY BECAME A KEY TRAINING FILTER
Generating plausible scientific prose was not sufficient for the project.
Early supervised fine-tuning improved overall report quality but still left AstaBrief behind Ai2's Claude-powered pipeline in answer precision and citation quality.
Ai2 tested several methods for filtering weaker synthetic training examples, including output-to-input token ratios, citation relevance, citation density and citation diversity.
Citation density produced the strongest improvement.
Reports containing substantial amounts of text without supporting citations were removed from the training data, giving the model a stronger statistical signal that substantive claims in scientific synthesis should remain connected to retrieved evidence.
More complicated combinations of filters did not produce corresponding improvements.
This result is relevant beyond AstaBrief because it demonstrates how specialization can come from the composition and filtering of post-training data, rather than requiring a substantially larger base model or additional scientific pretraining.
··········
ASTABRIEF USES SFT AND DPO INSTEAD OF REINFORCEMENT LEARNING
Ai2 considered reinforcement-learning approaches for improving long-form scientific reports but selected a simpler training pipeline.
AstaBrief first undergoes supervised fine-tuning, teaching the model the expected structure and evidence-grounded behavior using filtered report examples.
Ai2 then applies direct preference optimization, using pairs of candidate reports and preference judgments to reinforce outputs that better match the desired reporting behavior.
........
Training stage | Function |
Base model | Qwen3-8B provides the underlying language model |
Supervised fine-tuning | Teaches scientific report structure and evidence use |
Citation filtering | Removes weakly grounded training examples |
Direct preference optimization | Learns preferences between alternative reports |
Asta retrieval pipeline | Supplies scientific evidence during inference |
........
Ai2 chose this approach partly because reinforcement-learning pipelines can introduce additional computational cost and training instability.
The experiment therefore tests how far an 8B model can be specialized through carefully constructed post-training data and preference optimization before more complex training methods become necessary.
··········
THE MODEL IS COMPETITIVE WITH AI2'S MORE COMPLEX REPORT PIPELINES
Ai2 evaluated AstaBrief against its Claude-powered report pipeline and DR Tulu using measures covering report quality and citation behavior.
AstaBrief remained competitive across several automated measures despite using an 8B model and a substantially simpler inference architecture.
Ai2 also conducted a smaller human evaluation involving three scientific researchers and 14 research questions. Researchers compared reports on overall preference, completeness, relevance, organization and citation accuracy.
The evaluation did not establish AstaBrief as universally superior to the larger systems. DR Tulu performed better on overall preference, while two of the three researchers preferred AstaBrief on citation-accuracy measures.
Ai2 also notes an important limitation: much of the training and evaluation work was completed during 2025, and the proprietary models used as comparison points reflected the frontier available during development.
The results therefore evaluate the effectiveness of AstaBrief's specialization strategy rather than establish its position against every frontier model available in 2026.
··········
EARLY ASTA USAGE SHOWS FAST MODE IS ALREADY REPLACING THINKING MODE FOR SOME USERS
AstaBrief is already operating inside Asta rather than being released exclusively as an experimental model checkpoint.
Among the first 374 Asta users who tried Fast mode, Ai2 reports that 29.1% used it on at least two separate days.
Users generated an average of 3.67 report threads with the mode.
More significantly, 23% of users who tried Fast subsequently continued using it without returning to Thinking mode for later report threads. Another 18% alternated between the two modes depending on the task.
Positive-feedback rates were also close: 84.2% for Fast mode versus 85.2% for Thinking mode.
The sample remains relatively small and Ai2 explicitly treats the usage evidence as preliminary, but the results indicate that the faster pipeline can already satisfy a meaningful subset of scientific-report workflows without requiring the more compute-intensive alternative.
··········
OPEN WEIGHTS ALLOW SCIENTIFIC REPORTS TO RUN ON PRIVATE INFRASTRUCTURE
Ai2 is releasing AstaBrief's model weights rather than limiting access to the hosted Asta product.
Research institutions can therefore deploy the model on their own infrastructure and keep report generation behind organizational firewalls.
That property has practical relevance in scientific environments where even the research question itself can contain sensitive information about unpublished experiments, proprietary compounds, clinical research or work that has not yet been disclosed publicly.
Ai2 is also releasing an example workflow that allows researchers to generate reports from their own PDF collections.
AstaBrief therefore separates two components that are often bundled inside proprietary research assistants: the retrieval of scientific evidence and the model responsible for transforming retrieved evidence into the final synthesis.
Organizations can adapt the surrounding retrieval system while retaining the specialized report-generation model locally.
··········
AN 8B SPECIALIZED MODEL CHANGES THE COMPUTE PROFILE OF AI-ASSISTED RESEARCH
AstaBrief illustrates a different route from scaling a general-purpose model to hundreds of billions of parameters and applying it unchanged to scientific work.
The model begins with an 8B general-purpose foundation and concentrates additional training on report structure, evidence use, citations and scientific-query behavior.
The result is a system small enough to be distributed as open weights while performing a task that Ai2 previously handled through a more complicated pipeline backed by proprietary models.
The 3.5× pipeline-speed improvement does not demonstrate that AstaBrief is 3.5× faster than Claude as a language model. It reflects the combination of a smaller specialized model and an inference architecture that eliminates several intermediate stages.
That distinction is important when evaluating specialized AI systems: reducing the number of model calls and transformations required to complete a task can produce large efficiency gains without requiring equivalent improvements in raw model throughput.
··········
ASTABRIEF POINTS TOWARD SMALLER MODELS BUILT AROUND SPECIFIC SCIENTIFIC WORKFLOWS
Ai2 plans to continue developing the approach through stronger retrieval-plus-reinforcement-learning methods, multi-turn interactions, additional tools, new scientific data sources and query decomposition.
Future evaluation is also expected to move beyond checking whether a citation technically supports a sentence.
Scientific synthesis requires models to preserve evidentiary scope: a result observed in one population should not automatically become a universal conclusion, and a descriptive study should not be transformed into a recommendation simply because the generated prose sounds plausible.
AstaBrief addresses part of that problem through citation-focused training and retrieval-grounded generation.
Its broader significance lies in the architecture: a relatively small open model, specialized around a clearly defined scientific workflow, can replace several expensive generation stages while remaining deployable on infrastructure controlled by the researcher.
For scientific AI, that creates a development path based not only on larger models, but on better training examples, stronger evidence grounding and simpler task-specific inference pipelines.
··········
FOLLOW US FOR MORE.
DATA STUDIOS
datastudios.org




