Research

MachinoAI explainer / Retrieval-Augmented Generation

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Defines the canonical RAG architecture that combines a dense retriever, an external document index, and a seq2seq generator trained end-to-end.

Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, Douwe KielaMay 22, 2020NeurIPS 202012 min read
NEWRetrieval-Augmented GenerationAdvanced

01

Abstract

PAPER-DERIVED FACTS: RAG combines parametric memory in a pre-trained sequence-to-sequence model with non-parametric memory in a dense Wikipedia index. The paper introduces two variants that differ in whether one document or different documents can support generated tokens.

02

Introduction

PAPER-DERIVED FACTS: Large language models can store knowledge in parameters but are difficult to update and may hallucinate. RAG makes knowledge retrieval an explicit part of generation and enables the external memory to be replaced without retraining the generator.

03

Problem

PAPER-DERIVED FACTS: The goal is to perform knowledge-intensive NLP tasks while retaining flexible text generation. The model should retrieve relevant passages, condition generation on them, and learn retrieval and generation jointly from task input-output pairs.

04

Background

PAPER-DERIVED FACTS: REALM and DPR demonstrated learned retrieval for open-domain QA, while BART provided a strong pre-trained encoder-decoder generator. RAG combines these ideas into a general-purpose retrieval-conditioned generation framework.

05

Methodology

PAPER-DERIVED FACTS: RAG treats the retrieved document z as a latent variable and approximates marginalization over the top-k retrieved documents. RAG-Sequence keeps one document for the full output, whereas RAG-Token marginalizes a document distribution independently at each generated token.

Figure notes

Visual evidence

Figure 1
Retrieval-augmented generation pipeline — Retrieval-augmented generation pipelinehttps://arxiv.org/pdf/2005.11401

06

Architecture

PAPER-DERIVED FACTS: A DPR-style query encoder retrieves passages from a dense index using MIPS. The retrieved passage and query are fed to BART-large, and the retriever and generator are jointly optimized through the task likelihood.

07

Dataset

PAPER-DERIVED FACTS: The non-parametric memory is a December 2018 Wikipedia dump split into 100-word passages, yielding about 21 million documents. Evaluation covers Natural Questions, TriviaQA, WebQuestions, CuratedTREC, MS-MARCO, Jeopardy question generation, and FEVER.

08

Training

PAPER-DERIVED FACTS: The retriever is initialized from DPR and the generator from BART. Training minimizes negative marginal log-likelihood; the document encoder and index are kept fixed while the query encoder and generator are fine-tuned.

09

Experiments

PAPER-DERIVED FACTS: Experiments cover open-domain QA, abstractive QA, question generation, and fact verification. The paper also evaluates generation diversity, retrieval behavior, index replacement, and the effect of changing the number of retrieved documents.

10

Baselines

PAPER-DERIVED FACTS: Baselines include closed-book T5/BART-style generators, DPR-based extractive QA, REALM, and task-specific retrieval systems. The comparisons test whether generation plus retrieval can replace specialized extractive pipelines.

11

Results

PAPER-DERIVED FACTS: RAG-Sequence reaches 44.5 Exact Match on Natural Questions and reports strong results across several knowledge-intensive tasks. The paper also reports more factual, specific, and diverse generations than BART on the studied generation tasks.

12

Ablation

PAPER-DERIVED FACTS: The paper examines retrieval variants, the number of retrieved documents, retrieval index behavior, and the distinction between RAG-Sequence and RAG-Token. These experiments show that retrieval quality and the marginalization strategy materially affect results.

13

Limitations

Error Analysis

PAPER-DERIVED FACTS: The authors identify retrieval collapse on generation tasks with weak factual supervision, where the retriever may repeatedly select similar documents. RAG can also fail when the required evidence is absent from the indexed corpus.

14

Conclusion

PAPER-DERIVED FACTS: RAG establishes a modular architecture in which external knowledge can be retrieved at inference time and combined with a generative model. Its latent-document formulation became a basis for later retrieval-augmented systems.

15

References

PAPER-DERIVED FACTS: Primary source: Lewis, P. et al. (2020), Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020. arXiv:2005.11401.

Continue reading

Related research

Browse all research