MachinoAI explainer / Retrieval-Augmented Generation
REALM: Retrieval-Augmented Language Model Pre-Training
Introduces end-to-end retrieval during language-model pre-training so factual knowledge can live in an external, replaceable corpus rather than only in model parameters.
01
Abstract
PAPER-DERIVED FACTS: REALM augments language-model pre-training with a neural retriever over a large textual knowledge corpus. The paper jointly trains retrieval and prediction using a masked-language-model objective and then applies the learned retriever to open-domain question answering.
02
Introduction
PAPER-DERIVED FACTS: Parametric language models can memorize substantial world knowledge, but updating or inspecting that knowledge is difficult. REALM addresses this by adding a non-parametric memory that can be searched at pre-training, fine-tuning, and inference time.
03
Problem
PAPER-DERIVED FACTS: The target problem is knowledge-intensive prediction when relevant facts may be absent, stale, or difficult to expose in model parameters. The system must retrieve useful evidence while keeping the retrieval step trainable.
04
Background
PAPER-DERIVED FACTS: Earlier retrieval-augmented approaches such as ORQA combined language representations with learned retrieval, while standard BERT-style pre-training stored knowledge implicitly. REALM extends this direction to unsupervised pre-training with a large external corpus.
05
Methodology
PAPER-DERIVED FACTS: For an input x, REALM models p(y|x)=Σz p(y|z,x)p(z|x), treating the retrieved document z as a latent variable. A neural retriever scores candidate documents and a knowledge-augmented encoder predicts masked tokens or downstream answers.
Figure notes
Visual evidence
06
Architecture
PAPER-DERIVED FACTS: The system has a BERT-style input retriever, a textual knowledge corpus, and a knowledge-augmented encoder. Document retrieval uses dense representations and Maximum Inner Product Search (MIPS); retrieved text is joined with the input before encoding.
07
Dataset
PAPER-DERIVED FACTS: The retrieval corpus is based on Wikipedia. Pre-training uses unlabeled text with masked spans, while downstream evaluation focuses on open-domain question answering benchmarks including Natural Questions, WebQuestions, and CuratedTREC.
08
Training
PAPER-DERIVED FACTS: Pre-training uses masked language modeling as the learning signal, allowing gradients to update the retriever and encoder. The paper uses asynchronous MIPS index refreshes so document embeddings can be periodically rebuilt without stopping the main trainer.
09
Experiments
PAPER-DERIVED FACTS: Experiments compare REALM against models that store knowledge parametrically and systems with explicit retrieval. The key evaluation is open-domain QA after retrieval-aware pre-training and supervised fine-tuning.
10
Baselines
PAPER-DERIVED FACTS: The study compares with contemporary open-domain QA and language-model approaches, including BERT-style models, ORQA, and T5-style closed-book approaches. Comparisons cover both implicit and explicit knowledge storage.
11
Results
PAPER-DERIVED FACTS: REALM reports strong gains on open-domain QA and shows that retrieval-aware pre-training can outperform prior approaches while using substantially fewer parameters than very large closed-book models. The authors also report qualitative benefits from inspectable external evidence.
12
Ablation
PAPER-DERIVED FACTS: The paper studies components such as retrieval pre-training and retrieval behavior to establish that the gains are not obtained simply from adding an external index at inference time. The retrieval mechanism and its training signal are central to the reported improvements.
13
Limitations
Error Analysis
PAPER-DERIVED FACTS: Failure cases arise when retrieval does not surface the evidence needed by the encoder or when the relevant knowledge is difficult to connect to the masked prediction. The external index improves transparency but does not eliminate retrieval errors.
14
Conclusion
PAPER-DERIVED FACTS: REALM demonstrates that external textual memory can participate directly in language-model pre-training. Its central contribution is learning the retriever from the same prediction objective rather than treating retrieval as a fixed preprocessing component.
15
References
PAPER-DERIVED FACTS: Primary source: Guu, K., Lee, K., Tung, Z., Pasupat, P., & Chang, M.-W. (2020), REALM: Retrieval-Augmented Language Model Pre-Training. ICML 2020. arXiv:2002.08909.
Continue reading
Related research
Dense Passage Retrieval for Open-Domain Question Answering
Introduces a practical dual-encoder dense retriever that replaces sparse lexical matching with learned semantic passage representations.
Retrieval-Augmented GenerationFusion-in-Decoder
Introduces FiD, which scales multi-passage QA by encoding retrieved passages independently and fusing them inside the decoder.
Retrieval-Augmented GenerationRetrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Defines the canonical RAG architecture that combines a dense retriever, an external document index, and a seq2seq generator trained end-to-end.