MachinoAI explainer / Retrieval-Augmented Generation
Fusion-in-Decoder: Leveraging Passage Retrieval for Open-Domain Question Answering
Introduces FiD, which scales multi-passage QA by encoding retrieved passages independently and fusing them inside the decoder.
01
Abstract
PAPER-DERIVED FACTS: FiD processes each retrieved passage together with the question using a shared encoder and concatenates the resulting representations. The decoder attends over all encoded passages to generate the final answer.
02
Introduction
PAPER-DERIVED FACTS: Open-domain QA depends on retrieving evidence from a large corpus. A central design question is how to combine many retrieved passages without forcing the encoder to process their entire concatenation at once.
03
Problem
PAPER-DERIVED FACTS: The model must exploit information distributed across multiple retrieved passages while keeping encoding computationally manageable. Earlier approaches often fused documents too early or restricted the amount of evidence available to the generator.
04
Background
PAPER-DERIVED FACTS: FiD builds on encoder-decoder Transformers and dense passage retrieval. Its key design choice is to delay passage fusion until decoder cross-attention, allowing each passage to be encoded independently.
05
Methodology
PAPER-DERIVED FACTS: For each retrieved passage, FiD concatenates the question and passage and applies the same encoder. The encoded sequences are then concatenated and provided as the memory over which the decoder attends during generation.
Figure notes
Visual evidence
06
Architecture
PAPER-DERIVED FACTS: The architecture consists of a shared encoder applied independently to each question-passage pair and a decoder that performs cross-attention over the union of their encoded representations. Passage identity is preserved with special tokens.
07
Dataset
PAPER-DERIVED FACTS: Experiments use open-domain QA datasets including Natural Questions, TriviaQA, WebQuestions, and CuratedTREC. Retrieval candidates are produced by dense retrieval systems and passed to the generative reader.
08
Training
PAPER-DERIVED FACTS: FiD is trained end-to-end for answer generation using a sequence-to-sequence likelihood objective. The retriever can be fixed while the reader learns to exploit the retrieved evidence.
09
Experiments
PAPER-DERIVED FACTS: The paper evaluates the effect of the number of retrieved passages and compares FiD against extractive and generative QA systems. Scaling the number of passages tests whether additional evidence can be used effectively.
10
Baselines
PAPER-DERIVED FACTS: Comparisons include extractive readers, earlier retrieval-augmented generative readers, and T5-based approaches. Retrieval quality is held as a separate component from the reader architecture.
11
Results
PAPER-DERIVED FACTS: FiD achieves strong open-domain QA performance and shows that increasing the number of retrieved passages can continue to improve answer quality. Its decoder-level fusion is particularly effective when evidence is distributed across documents.
12
Ablation
PAPER-DERIVED FACTS: The study varies the number of retrieved passages and model scale to examine how much evidence the decoder can use. Results show a clear relationship between available retrieved context and QA performance.
13
Limitations
Error Analysis
PAPER-DERIVED FACTS: FiD remains dependent on retrieval quality: irrelevant or missing passages constrain the decoder. More passages also increase computation, so retrieval breadth introduces a latency and memory trade-off.
14
Conclusion
PAPER-DERIVED FACTS: FiD demonstrates that late fusion is an effective way to combine many retrieved passages in generative QA. The architecture separates passage encoding from evidence aggregation and became influential in retrieval-augmented readers.
15
References
PAPER-DERIVED FACTS: Primary source: Izacard, G. & Grave, E. (2021), Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering / Fusion-in-Decoder. arXiv:2007.01282.
Continue reading
Related research
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Defines the canonical RAG architecture that combines a dense retriever, an external document index, and a seq2seq generator trained end-to-end.
Retrieval-Augmented GenerationDense Passage Retrieval for Open-Domain Question Answering
Introduces a practical dual-encoder dense retriever that replaces sparse lexical matching with learned semantic passage representations.
Retrieval-Augmented GenerationREALM
Introduces end-to-end retrieval during language-model pre-training so factual knowledge can live in an external, replaceable corpus rather than only in model parameters.