Research

MachinoAI explainer / Retrieval-Augmented Generation

REALM: Retrieval-Augmented Language Model Pre-Training

Introduces end-to-end retrieval during language-model pre-training so factual knowledge can live in an external, replaceable corpus rather than only in model parameters.

Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, Ming-Wei ChangFeb 10, 2020ICML 202012 min read
NEWRetrieval-Augmented GenerationAdvanced

01

Abstract

PAPER-DERIVED FACTS: REALM augments language-model pre-training with a neural retriever over a large textual knowledge corpus. The paper jointly trains retrieval and prediction using a masked-language-model objective and then applies the learned retriever to open-domain question answering.

02

Introduction

PAPER-DERIVED FACTS: Parametric language models can memorize substantial world knowledge, but updating or inspecting that knowledge is difficult. REALM addresses this by adding a non-parametric memory that can be searched at pre-training, fine-tuning, and inference time.

03

Problem

PAPER-DERIVED FACTS: The target problem is knowledge-intensive prediction when relevant facts may be absent, stale, or difficult to expose in model parameters. The system must retrieve useful evidence while keeping the retrieval step trainable.

04

Background

PAPER-DERIVED FACTS: Earlier retrieval-augmented approaches such as ORQA combined language representations with learned retrieval, while standard BERT-style pre-training stored knowledge implicitly. REALM extends this direction to unsupervised pre-training with a large external corpus.

05

Methodology

PAPER-DERIVED FACTS: For an input x, REALM models p(y|x)=Σz p(y|z,x)p(z|x), treating the retrieved document z as a latent variable. A neural retriever scores candidate documents and a knowledge-augmented encoder predicts masked tokens or downstream answers.

Figure notes

Visual evidence

Figure 1
REALM retrieval-augmented pre-training workflow — REALM retrieval-augmented pre-training workflowhttps://arxiv.org/pdf/2002.08909

06

Architecture

PAPER-DERIVED FACTS: The system has a BERT-style input retriever, a textual knowledge corpus, and a knowledge-augmented encoder. Document retrieval uses dense representations and Maximum Inner Product Search (MIPS); retrieved text is joined with the input before encoding.

07

Dataset

PAPER-DERIVED FACTS: The retrieval corpus is based on Wikipedia. Pre-training uses unlabeled text with masked spans, while downstream evaluation focuses on open-domain question answering benchmarks including Natural Questions, WebQuestions, and CuratedTREC.

08

Training

PAPER-DERIVED FACTS: Pre-training uses masked language modeling as the learning signal, allowing gradients to update the retriever and encoder. The paper uses asynchronous MIPS index refreshes so document embeddings can be periodically rebuilt without stopping the main trainer.

09

Experiments

PAPER-DERIVED FACTS: Experiments compare REALM against models that store knowledge parametrically and systems with explicit retrieval. The key evaluation is open-domain QA after retrieval-aware pre-training and supervised fine-tuning.

10

Baselines

PAPER-DERIVED FACTS: The study compares with contemporary open-domain QA and language-model approaches, including BERT-style models, ORQA, and T5-style closed-book approaches. Comparisons cover both implicit and explicit knowledge storage.

11

Results

PAPER-DERIVED FACTS: REALM reports strong gains on open-domain QA and shows that retrieval-aware pre-training can outperform prior approaches while using substantially fewer parameters than very large closed-book models. The authors also report qualitative benefits from inspectable external evidence.

12

Ablation

PAPER-DERIVED FACTS: The paper studies components such as retrieval pre-training and retrieval behavior to establish that the gains are not obtained simply from adding an external index at inference time. The retrieval mechanism and its training signal are central to the reported improvements.

13

Limitations

Error Analysis

PAPER-DERIVED FACTS: Failure cases arise when retrieval does not surface the evidence needed by the encoder or when the relevant knowledge is difficult to connect to the masked prediction. The external index improves transparency but does not eliminate retrieval errors.

14

Conclusion

PAPER-DERIVED FACTS: REALM demonstrates that external textual memory can participate directly in language-model pre-training. Its central contribution is learning the retriever from the same prediction objective rather than treating retrieval as a fixed preprocessing component.

15

References

PAPER-DERIVED FACTS: Primary source: Guu, K., Lee, K., Tung, Z., Pasupat, P., & Chang, M.-W. (2020), REALM: Retrieval-Augmented Language Model Pre-Training. ICML 2020. arXiv:2002.08909.

Continue reading

Related research

Browse all research