Research

MachinoAI explainer / Retrieval-Augmented Generation

Improving Language Models by Retrieving from Trillions of Tokens

Shows that retrieval from a 2-trillion-token database can substantially improve language modeling while using far fewer model parameters than similarly capable dense models.

Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, Laurent SifreDec 8, 2021arXiv 202112 min read
NEWRetrieval-Augmented GenerationAdvanced

01

Abstract

PAPER-DERIVED FACTS: RETRO conditions autoregressive language modeling on text chunks retrieved from a very large database. The reported system uses a 2-trillion-token database and achieves performance comparable to much larger dense language models with substantially fewer parameters.

02

Introduction

PAPER-DERIVED FACTS: Scaling Transformer parameter counts improves language modeling but increases compute and memory costs. RETRO explores a complementary axis: give a smaller model access to a much larger external text memory through retrieval.

03

Problem

PAPER-DERIVED FACTS: The problem is how to incorporate massive external corpora into autoregressive generation without requiring all knowledge to be compressed into parameters. Retrieval must be efficient and its representations must be integrated into each Transformer block.

04

Background

PAPER-DERIVED FACTS: Earlier retrieval-augmented models often focused on question answering or modest-scale document indexes. RETRO extends retrieval conditioning directly to general language modeling and to a database containing trillions of tokens.

05

Methodology

PAPER-DERIVED FACTS: The model retrieves nearest-neighbor chunks based on preceding text, encodes the retrieved neighbors, and uses chunked cross-attention to condition Transformer representations. The retriever is frozen while the language model learns to exploit retrieved memory.

Figure notes

Visual evidence

Figure 1
RETRO block with retrieval and cross-attention — RETRO block with retrieval and cross-attentionhttps://arxiv.org/pdf/2112.04426

06

Architecture

PAPER-DERIVED FACTS: RETRO uses a standard autoregressive Transformer augmented with a retrieval mechanism. Retrieved chunks are encoded by a separate Transformer encoder and injected through cross-attention inside RETRO blocks, separating retrieval from ordinary self-attention.

07

Dataset

PAPER-DERIVED FACTS: The retrieval database contains roughly 2 trillion tokens assembled from web pages, books, news, and code. Training and evaluation use large-scale language-model corpora including The Pile for reported comparisons.

08

Training

PAPER-DERIVED FACTS: RETRO models are trained from scratch with retrieval conditioning, while the retrieval system can also be attached to existing pretrained Transformers. The training objective remains autoregressive next-token prediction.

09

Experiments

PAPER-DERIVED FACTS: The study compares RETRO with dense language models at different parameter scales and evaluates both perplexity and downstream knowledge-intensive tasks such as question answering. It also examines retrieval database size and model scaling.

10

Baselines

PAPER-DERIVED FACTS: Key comparisons include GPT-3 and Jurassic-1 style dense language models, as well as dense Transformer baselines with substantially larger parameter counts. The central comparison is parameter efficiency versus external-memory access.

11

Results

PAPER-DERIVED FACTS: The 7.5B-parameter RETRO model is reported to achieve language-modeling performance comparable to much larger models, including GPT-3 and Jurassic-1, while using a 2-trillion-token retrieval database. Fine-tuning also transfers the benefits to downstream QA.

12

Ablation

PAPER-DERIVED FACTS: The work studies retrieval database scale, retrieval frequency, and the contribution of retrieved context to model performance. The results show that more external memory can provide useful gains without proportionally increasing model parameters.

13

Limitations

Error Analysis

PAPER-DERIVED FACTS: Retrieval quality and coverage constrain the benefit of external memory. Retrieved chunks can be irrelevant, redundant, or poorly aligned with the current generation context, and retrieval also adds systems complexity and latency.

14

Conclusion

PAPER-DERIVED FACTS: RETRO establishes retrieval-conditioned language modeling at unprecedented corpus scale. The paper argues that model capacity and external memory can be traded against one another rather than relying solely on parameter scaling.

15

References

PAPER-DERIVED FACTS: Primary source: Borgeaud, S. et al. (2022), Improving Language Models by Retrieving from Trillions of Tokens. arXiv:2112.04426.

Continue reading

Related research

Browse all research