Research

MachinoAI explainer / Retrieval-Augmented Generation

Dense Passage Retrieval for Open-Domain Question Answering

Introduces a practical dual-encoder dense retriever that replaces sparse lexical matching with learned semantic passage representations.

Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau YihApr 10, 2020EMNLP 202012 min read
NEWRetrieval-Augmented GenerationAdvanced

01

Abstract

PAPER-DERIVED FACTS: DPR learns dense vector representations for questions and passages using two BERT encoders. Retrieval is performed with inner-product similarity and the resulting passages can be passed to a downstream reader.

02

Introduction

PAPER-DERIVED FACTS: Traditional open-domain QA systems commonly relied on sparse retrieval such as BM25. DPR investigates whether a small amount of supervised QA data is enough to learn a semantic retriever that is both effective and computationally practical.

03

Problem

PAPER-DERIVED FACTS: The task is to retrieve passages containing answers from a large corpus given a natural-language question. The challenge is to learn semantic relevance while keeping retrieval fast enough for millions of passages.

04

Background

PAPER-DERIVED FACTS: Sparse methods score lexical overlap, while dense retrieval maps queries and passages into continuous vector spaces. Bi-encoder architectures allow passage embeddings to be precomputed and indexed for efficient approximate nearest-neighbor search.

05

Methodology

PAPER-DERIVED FACTS: DPR encodes a question q and passage p independently and scores them using the dot product qᵀp. Training uses positive passages and negative passages, including hard negatives mined from strong traditional retrievers.

Figure notes

Visual evidence

Figure 1
DPR retriever-reader architecture — DPR retriever-reader architecturehttps://arxiv.org/pdf/2004.04906

06

Architecture

PAPER-DERIVED FACTS: The retriever contains separate question and passage BERT encoders. Passage vectors are precomputed and stored in an index; at query time the question vector is compared against passage vectors and the top-k candidates are returned.

07

Dataset

PAPER-DERIVED FACTS: Training uses open-domain QA datasets with question-answer pairs and supporting passages. Evaluation covers Natural Questions, TriviaQA, WebQuestions, CuratedTREC, and other open-domain retrieval/QA benchmarks.

08

Training

PAPER-DERIVED FACTS: The dual encoders are optimized with a contrastive objective that increases the score of positive passages relative to negatives. Hard-negative passages are important because random negatives are often too easy.

09

Experiments

PAPER-DERIVED FACTS: The study evaluates retrieval recall and end-to-end QA after passing retrieved passages to a reader. It compares DPR against BM25 and earlier retrieval systems across multiple QA datasets.

10

Baselines

PAPER-DERIVED FACTS: The principal baseline is Lucene BM25, along with prior learned and hybrid retrieval systems. End-to-end comparisons use readers operating on retrieved candidate passages.

11

Results

PAPER-DERIVED FACTS: DPR improves top-20 passage retrieval accuracy over a strong BM25 baseline by large margins on the reported datasets, and the resulting reader systems achieve new state-of-the-art results on several open-domain QA benchmarks.

12

Ablation

PAPER-DERIVED FACTS: The paper studies negative sampling, hard negatives, training-data size, and encoder configurations. The results show that supervised semantic learning and informative negatives are key to dense retrieval quality.

13

Limitations

Error Analysis

PAPER-DERIVED FACTS: Dense retrieval can miss passages whose relevance depends on exact lexical cues or rare entities, and retrieval quality remains sensitive to the training distribution. Errors at the retriever stage propagate to the downstream reader.

14

Conclusion

PAPER-DERIVED FACTS: DPR demonstrates that simple dense dual encoders can provide strong open-domain passage retrieval at scale. Its design became a core building block for subsequent RAG architectures.

15

References

PAPER-DERIVED FACTS: Primary source: Karpukhin, V. et al. (2020), Dense Passage Retrieval for Open-Domain Question Answering. EMNLP 2020. arXiv:2004.04906.

Continue reading

Related research

Browse all research