MachinoAI explainer / Retrieval-Augmented Generation
Active Retrieval Augmented Generation
Introduces active retrieval that dynamically decides when and what to retrieve while generating long-form answers.
01
Abstract
PAPER-DERIVED FACTS: FLARE addresses long-form generation where a single retrieval step may not provide evidence for later sentences. It actively predicts upcoming content, retrieves evidence when confidence is low, and regenerates using the retrieved information.
02
Introduction
PAPER-DERIVED FACTS: Standard RAG usually retrieves once before generation. For long responses, the information need changes as generation progresses, motivating a retrieval policy that can react to intermediate model uncertainty.
03
Problem
PAPER-DERIVED FACTS: The system must decide both when retrieval is necessary and what query should be issued during generation. Retrieval should be triggered selectively rather than on every token or sentence.
04
Background
PAPER-DERIVED FACTS: FLARE builds on retrieval-augmented language generation and uncertainty-aware generation. Its core idea is to use a draft of future content as a signal for targeted retrieval.
05
Methodology
PAPER-DERIVED FACTS: FLARE generates a temporary future sentence or segment, checks token probabilities against a confidence threshold, and triggers retrieval when the draft is uncertain. Retrieved passages are then used to regenerate the segment with external evidence.
Figure notes
Visual evidence
06
Architecture
PAPER-DERIVED FACTS: The loop consists of generation, confidence checking, query construction, retrieval, and evidence-grounded regeneration. The process repeats until the response is complete.
07
Dataset
PAPER-DERIVED FACTS: Experiments cover long-form knowledge-intensive generation tasks and datasets where information needs evolve across a response. The method is evaluated against static retrieval baselines.
08
Training
PAPER-DERIVED FACTS: FLARE is primarily a retrieval-and-generation strategy that can be applied to existing language models. It does not require retraining a large generator for the active retrieval policy described in the paper.
09
Experiments
PAPER-DERIVED FACTS: The paper compares active retrieval against retrieve-once approaches and evaluates factuality and generation quality across long-form tasks. Different confidence thresholds control retrieval frequency.
10
Baselines
PAPER-DERIVED FACTS: Baselines include standard retrieval-augmented generation that retrieves from the initial query and other retrieval-based generation approaches. Comparisons focus on whether dynamic retrieval improves factual long-form output.
11
Results
PAPER-DERIVED FACTS: FLARE improves factuality and generation quality on the studied long-form tasks while avoiding unnecessary retrieval when the model is already confident.
12
Ablation
PAPER-DERIVED FACTS: The study varies retrieval thresholds and retrieval timing to show the trade-off between evidence coverage and retrieval cost. Too aggressive retrieval increases overhead, while weak thresholds can miss useful evidence.
13
Limitations
Error Analysis
PAPER-DERIVED FACTS: FLARE can still fail when the generated draft produces a poor retrieval query or when the external corpus lacks the needed evidence. Confidence estimates are also imperfect signals for factual correctness.
14
Conclusion
PAPER-DERIVED FACTS: Active retrieval makes RAG adaptive to changing information needs during generation. The paper provides an important template for iterative and agentic retrieval loops.
15
References
PAPER-DERIVED FACTS: Primary source: Jiang, Z. et al. (2023), Active Retrieval Augmented Generation. EMNLP 2023. arXiv:2305.06983.
Continue reading
Related research
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Defines the canonical RAG architecture that combines a dense retriever, an external document index, and a seq2seq generator trained end-to-end.
Retrieval-Augmented GenerationDense Passage Retrieval for Open-Domain Question Answering
Introduces a practical dual-encoder dense retriever that replaces sparse lexical matching with learned semantic passage representations.
Retrieval-Augmented GenerationREALM
Introduces end-to-end retrieval during language-model pre-training so factual knowledge can live in an external, replaceable corpus rather than only in model parameters.