MachinoAI explainer / Retrieval-Augmented Generation
Improving Language Models by Retrieving from Trillions of Tokens
Shows that retrieval from a 2-trillion-token database can substantially improve language modeling while using far fewer model parameters than similarly capable dense models.
01
Abstract
PAPER-DERIVED FACTS: RETRO conditions autoregressive language modeling on text chunks retrieved from a very large database. The reported system uses a 2-trillion-token database and achieves performance comparable to much larger dense language models with substantially fewer parameters.
02
Introduction
PAPER-DERIVED FACTS: Scaling Transformer parameter counts improves language modeling but increases compute and memory costs. RETRO explores a complementary axis: give a smaller model access to a much larger external text memory through retrieval.
03
Problem
PAPER-DERIVED FACTS: The problem is how to incorporate massive external corpora into autoregressive generation without requiring all knowledge to be compressed into parameters. Retrieval must be efficient and its representations must be integrated into each Transformer block.
04
Background
PAPER-DERIVED FACTS: Earlier retrieval-augmented models often focused on question answering or modest-scale document indexes. RETRO extends retrieval conditioning directly to general language modeling and to a database containing trillions of tokens.
05
Methodology
PAPER-DERIVED FACTS: The model retrieves nearest-neighbor chunks based on preceding text, encodes the retrieved neighbors, and uses chunked cross-attention to condition Transformer representations. The retriever is frozen while the language model learns to exploit retrieved memory.
Figure notes
Visual evidence
06
Architecture
PAPER-DERIVED FACTS: RETRO uses a standard autoregressive Transformer augmented with a retrieval mechanism. Retrieved chunks are encoded by a separate Transformer encoder and injected through cross-attention inside RETRO blocks, separating retrieval from ordinary self-attention.
07
Dataset
PAPER-DERIVED FACTS: The retrieval database contains roughly 2 trillion tokens assembled from web pages, books, news, and code. Training and evaluation use large-scale language-model corpora including The Pile for reported comparisons.
08
Training
PAPER-DERIVED FACTS: RETRO models are trained from scratch with retrieval conditioning, while the retrieval system can also be attached to existing pretrained Transformers. The training objective remains autoregressive next-token prediction.
09
Experiments
PAPER-DERIVED FACTS: The study compares RETRO with dense language models at different parameter scales and evaluates both perplexity and downstream knowledge-intensive tasks such as question answering. It also examines retrieval database size and model scaling.
10
Baselines
PAPER-DERIVED FACTS: Key comparisons include GPT-3 and Jurassic-1 style dense language models, as well as dense Transformer baselines with substantially larger parameter counts. The central comparison is parameter efficiency versus external-memory access.
11
Results
PAPER-DERIVED FACTS: The 7.5B-parameter RETRO model is reported to achieve language-modeling performance comparable to much larger models, including GPT-3 and Jurassic-1, while using a 2-trillion-token retrieval database. Fine-tuning also transfers the benefits to downstream QA.
12
Ablation
PAPER-DERIVED FACTS: The work studies retrieval database scale, retrieval frequency, and the contribution of retrieved context to model performance. The results show that more external memory can provide useful gains without proportionally increasing model parameters.
13
Limitations
Error Analysis
PAPER-DERIVED FACTS: Retrieval quality and coverage constrain the benefit of external memory. Retrieved chunks can be irrelevant, redundant, or poorly aligned with the current generation context, and retrieval also adds systems complexity and latency.
14
Conclusion
PAPER-DERIVED FACTS: RETRO establishes retrieval-conditioned language modeling at unprecedented corpus scale. The paper argues that model capacity and external memory can be traded against one another rather than relying solely on parameter scaling.
15
References
PAPER-DERIVED FACTS: Primary source: Borgeaud, S. et al. (2022), Improving Language Models by Retrieving from Trillions of Tokens. arXiv:2112.04426.
Continue reading
Related research
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Defines the canonical RAG architecture that combines a dense retriever, an external document index, and a seq2seq generator trained end-to-end.
Retrieval-Augmented GenerationDense Passage Retrieval for Open-Domain Question Answering
Introduces a practical dual-encoder dense retriever that replaces sparse lexical matching with learned semantic passage representations.
Retrieval-Augmented GenerationREALM
Introduces end-to-end retrieval during language-model pre-training so factual knowledge can live in an external, replaceable corpus rather than only in model parameters.