Research

MachinoAI explainer / Retrieval-Augmented Generation

Atlas: Few-shot Learning with Retrieval Augmented Language Models

Shows that retrieval-augmented language models can learn knowledge-intensive tasks effectively in few-shot settings.

Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, Edouard GraveAug 5, 2022JMLR / arXiv 202212 min read
NEWRetrieval-Augmented GenerationAdvanced

01

Abstract

PAPER-DERIVED FACTS: Atlas studies retrieval-augmented language modeling when only a few labeled examples are available. The model combines a learned retriever with a generative reader and is pre-trained to use external knowledge effectively.

02

Introduction

PAPER-DERIVED FACTS: Few-shot learning is difficult for knowledge-intensive tasks because a model must infer both the task and the relevant world knowledge from limited supervision. Atlas uses retrieval to expose external evidence instead of requiring all knowledge to be stored parametrically.

03

Problem

PAPER-DERIVED FACTS: The paper asks whether retrieval-augmented models can remain effective when downstream supervision is extremely small. The system must learn retrieval, generation, and task behavior from limited examples.

04

Background

PAPER-DERIVED FACTS: Atlas builds on dense retrieval, retrieval-augmented generation, and Fusion-in-Decoder. It treats an external text collection as non-parametric memory and uses a generative model to synthesize answers from retrieved passages.

05

Methodology

PAPER-DERIVED FACTS: Atlas jointly optimizes a retriever and a language model while using retrieved passages as context. The model is pre-trained on multiple objectives and then adapted to few-shot downstream tasks.

Figure notes

Visual evidence

Figure 1
Atlas retrieval-augmented language model — Atlas retrieval-augmented language modelhttps://arxiv.org/pdf/2208.03299

06

Architecture

PAPER-DERIVED FACTS: The architecture uses a retriever to select passages from a large index and a T5-style encoder-decoder to generate outputs. Retrieved passages are encoded separately and fused in the decoder in the FiD style.

07

Dataset

PAPER-DERIVED FACTS: Pre-training uses large text corpora and retrieval collections, while evaluation covers knowledge-intensive tasks such as open-domain QA and fact checking. The study explicitly tests very small labeled training sets.

08

Training

PAPER-DERIVED FACTS: The model is trained to retrieve useful passages and generate targets conditioned on them. Few-shot fine-tuning then uses only a small number of examples for the downstream task.

09

Experiments

PAPER-DERIVED FACTS: Experiments evaluate few-shot performance across QA, fact checking, and other knowledge-intensive benchmarks. Comparisons include retrieval-augmented and closed-book language models at different parameter scales.

10

Baselines

PAPER-DERIVED FACTS: Baselines include closed-book T5 models, earlier retrieval-augmented systems, and models using substantially more parameters. The goal is to isolate the benefit of external retrieval under limited supervision.

11

Results

PAPER-DERIVED FACTS: Atlas reports strong few-shot results and shows that retrieval augmentation can make relatively compact models competitive with much larger closed-book systems on knowledge-intensive tasks.

12

Ablation

PAPER-DERIVED FACTS: The work studies retriever quality, model scale, number of retrieved passages, and pre-training choices. The experiments show that both retrieval and the ability to use retrieved evidence are necessary for robust few-shot gains.

13

Limitations

Error Analysis

PAPER-DERIVED FACTS: Errors can result from retrieving passages that are topically related but insufficient for the exact question, or from the generator failing to synthesize multiple pieces of evidence. Few-shot settings amplify sensitivity to retrieval quality.

14

Conclusion

PAPER-DERIVED FACTS: Atlas demonstrates that retrieval can serve as an important source of transferable knowledge for few-shot learning. It provides a bridge between large-scale pre-training and low-data downstream adaptation.

15

References

PAPER-DERIVED FACTS: Primary source: Izacard, G. et al. (2022), Atlas: Few-shot Learning with Retrieval Augmented Language Models. arXiv:2208.03299.

Continue reading

Related research

Browse all research