Ai2 Open-Sources AstaBrief 8B, a Cited-Report Model for Research Questions
Ai2 has released AstaBrief 8B, an 8B open model that turns a research question and literature excerpts into a cited report. It powers Fast mode in Asta's report
On October 2, 2026, the Allen Institute for AI (Ai2) announced AstaBrief 8B on its official blog: an open model that turns a research question plus retrieved literature excerpts into a report with citations. The release ships open weights and training data so teams can run it on their own hardware and inspect or build on the approach.

Image source: Ai2 (@allen_ai)
The launch is tied to Asta, Ai2's platform for answering research questions from the scientific literature. AstaBrief 8B powers Fast mode in Asta's "Generate a report" feature, offered alongside the Claude-powered Thinking mode.
Fast mode vs. Thinking mode: one pass against a staged pipeline
Fast mode, driven by AstaBrief, skips the intermediate steps that summarize and group the retrieved excerpts and writes the report in one pass. According to Ai2, the design goal is to get researchers a preliminary report sooner.
The speed figures need their scope stated plainly. Ai2 reports an average of 51.1 seconds per report for Fast mode versus 178.5 seconds for Thinking mode — about 3.5x faster — measured end to end across Asta's full pipeline, not as standalone model inference latency.
On quality, Ai2 says AstaBrief is competitive with Thinking mode and DR Tulu on answer and citation measures in its own evaluations. That claim rests on Ai2's internal evaluation criteria; independent third-party reproduction has not yet been confirmed.
Built on Qwen3-8B, with data and a local-PDF example
The model is built on Qwen3-8B. Weights are published at allenai/AstaBrief_8B on Hugging Face, alongside the training data and an example repository for generating reports from local PDFs.
Ai2 says training queries came from ScholarQA and its earlier literature-synthesis systems. After removing bot and test traffic plus very short prompts, and filtering out non-English, non-scientific, or personal-information-containing requests with a language model, 90K queries remained. Ai2 generated cited reports from those queries, then quality-filtered them down to 39.5K examples for supervised fine-tuning (SFT). Using separate queries, it built report pairs where two judge models agreed — about 6K pairs — and trained with direct preference optimization (DPO).
Practical meaning and confirmed limits
For researchers and teams that already hold the evidence (retrieved excerpts), AstaBrief works as a writing layer that drafts a cited report quickly. Ai2 also frames open weights as the route for institutions to run the model on their own infrastructure when research questions touch sensitive or unpublished work.
Limits are confirmed: AstaBrief itself does not replace search, retrieval, or PDF ingestion — it expects a research question and retrieved literature excerpts as input. The latency figures are full-pipeline end-to-end numbers, and quality-superiority claims are confined to Ai2's own evaluations.
Sources
- Ai2 official blog: Open-sourcing AstaBrief, the fast report-generation model in Asta
- Ai2 official X (@allen_ai): AstaBrief 8B announcement thread
- Hugging Face: allenai/AstaBrief_8B