MachinoAI explainer / Retrieval-Augmented Generation
Atlas: Few-shot Learning with Retrieval Augmented Language Models
Shows that retrieval-augmented language models can learn knowledge-intensive tasks effectively in few-shot settings.
01
Abstract
PAPER-DERIVED FACTS: Atlas studies retrieval-augmented language modeling when only a few labeled examples are available. The model combines a learned retriever with a generative reader and is pre-trained to use external knowledge effectively.
02
Introduction
PAPER-DERIVED FACTS: Few-shot learning is difficult for knowledge-intensive tasks because a model must infer both the task and the relevant world knowledge from limited supervision. Atlas uses retrieval to expose external evidence instead of requiring all knowledge to be stored parametrically.
03
Problem
PAPER-DERIVED FACTS: The paper asks whether retrieval-augmented models can remain effective when downstream supervision is extremely small. The system must learn retrieval, generation, and task behavior from limited examples.
04
Background
PAPER-DERIVED FACTS: Atlas builds on dense retrieval, retrieval-augmented generation, and Fusion-in-Decoder. It treats an external text collection as non-parametric memory and uses a generative model to synthesize answers from retrieved passages.
05
Methodology
PAPER-DERIVED FACTS: Atlas jointly optimizes a retriever and a language model while using retrieved passages as context. The model is pre-trained on multiple objectives and then adapted to few-shot downstream tasks.
Figure notes
Visual evidence
06
Architecture
PAPER-DERIVED FACTS: The architecture uses a retriever to select passages from a large index and a T5-style encoder-decoder to generate outputs. Retrieved passages are encoded separately and fused in the decoder in the FiD style.
07
Dataset
PAPER-DERIVED FACTS: Pre-training uses large text corpora and retrieval collections, while evaluation covers knowledge-intensive tasks such as open-domain QA and fact checking. The study explicitly tests very small labeled training sets.
08
Training
PAPER-DERIVED FACTS: The model is trained to retrieve useful passages and generate targets conditioned on them. Few-shot fine-tuning then uses only a small number of examples for the downstream task.
09
Experiments
PAPER-DERIVED FACTS: Experiments evaluate few-shot performance across QA, fact checking, and other knowledge-intensive benchmarks. Comparisons include retrieval-augmented and closed-book language models at different parameter scales.
10
Baselines
PAPER-DERIVED FACTS: Baselines include closed-book T5 models, earlier retrieval-augmented systems, and models using substantially more parameters. The goal is to isolate the benefit of external retrieval under limited supervision.
11
Results
PAPER-DERIVED FACTS: Atlas reports strong few-shot results and shows that retrieval augmentation can make relatively compact models competitive with much larger closed-book systems on knowledge-intensive tasks.
12
Ablation
PAPER-DERIVED FACTS: The work studies retriever quality, model scale, number of retrieved passages, and pre-training choices. The experiments show that both retrieval and the ability to use retrieved evidence are necessary for robust few-shot gains.
13
Limitations
Error Analysis
PAPER-DERIVED FACTS: Errors can result from retrieving passages that are topically related but insufficient for the exact question, or from the generator failing to synthesize multiple pieces of evidence. Few-shot settings amplify sensitivity to retrieval quality.
14
Conclusion
PAPER-DERIVED FACTS: Atlas demonstrates that retrieval can serve as an important source of transferable knowledge for few-shot learning. It provides a bridge between large-scale pre-training and low-data downstream adaptation.
15
References
PAPER-DERIVED FACTS: Primary source: Izacard, G. et al. (2022), Atlas: Few-shot Learning with Retrieval Augmented Language Models. arXiv:2208.03299.
Continue reading
Related research
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Defines the canonical RAG architecture that combines a dense retriever, an external document index, and a seq2seq generator trained end-to-end.
Retrieval-Augmented GenerationDense Passage Retrieval for Open-Domain Question Answering
Introduces a practical dual-encoder dense retriever that replaces sparse lexical matching with learned semantic passage representations.
Retrieval-Augmented GenerationREALM
Introduces end-to-end retrieval during language-model pre-training so factual knowledge can live in an external, replaceable corpus rather than only in model parameters.