Research

MachinoAI explainer / Retrieval-Augmented Generation

Precise Zero-Shot Dense Retrieval without Relevance Labels

Introduces Hypothetical Document Embeddings (HyDE), a query-transformation technique for zero-shot dense retrieval without relevance labels.

Luyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanDec 20, 2022ACL 202312 min read
NEWRetrieval-Augmented GenerationAdvanced

01

Abstract

PAPER-DERIVED FACTS: HyDE addresses zero-shot dense retrieval when relevance labels are unavailable. An instruction-following language model generates a hypothetical document, an unsupervised encoder embeds it, and the resulting vector retrieves real documents from the target corpus.

02

Introduction

PAPER-DERIVED FACTS: Dense retrievers are usually strongest after task-specific supervision, but many domains lack labeled query-document pairs. HyDE asks whether generated text can provide a better semantic bridge between a query and an unseen corpus.

03

Problem

PAPER-DERIVED FACTS: A short user query may be a poor match for document embeddings because query and document language distributions differ. The challenge is to construct a retrieval representation without relevance supervision.

04

Background

PAPER-DERIVED FACTS: HyDE builds on dense retrieval and unsupervised encoders such as Contriever. It differs from direct query embedding by introducing an intermediate generated document that resembles the target corpus.

05

Methodology

PAPER-DERIVED FACTS: Given a query, an instruction-following LLM generates a hypothetical answer or document. An unsupervised contrastive encoder embeds that text, and nearest-neighbor search retrieves actual corpus documents using the generated embedding.

Figure notes

Visual evidence

Figure 1
HyDE hypothetical document retrieval workflow — HyDE hypothetical document retrieval workflowhttps://arxiv.org/pdf/2212.10496

06

Architecture

PAPER-DERIVED FACTS: The pipeline is query → hypothetical document generation → document embedding → dense retrieval. The LLM is not used as the final source of truth; the real corpus remains the evidence source.

07

Dataset

PAPER-DERIVED FACTS: Experiments span multiple retrieval settings, languages, and tasks including web search, question answering, and fact verification. The method is designed specifically for settings without relevance labels.

08

Training

PAPER-DERIVED FACTS: The generator can be an instruction-following model used without task-specific retrieval training. The dense encoder is pre-trained independently and used as an unsupervised retrieval model.

09

Experiments

PAPER-DERIVED FACTS: HyDE is evaluated against direct dense retrieval and supervised retrievers across several zero-shot retrieval benchmarks. The paper examines both English and multilingual settings.

10

Baselines

PAPER-DERIVED FACTS: The main unsupervised baseline is Contriever, while comparisons also include supervised dense retrievers where available. HyDE is evaluated as a retrieval transformation rather than a new encoder-training objective.

11

Results

PAPER-DERIVED FACTS: HyDE substantially improves over the unsupervised Contriever baseline and reaches performance comparable to strong fine-tuned retrievers on several tasks reported by the authors.

12

Ablation

PAPER-DERIVED FACTS: The study examines the role of hypothetical-document generation and the choice of encoder. Results support the hypothesis that generated text can bridge the semantic mismatch between short queries and full documents.

13

Limitations

Error Analysis

PAPER-DERIVED FACTS: Hypothetical documents can hallucinate details, but those details are not directly returned to the user; they influence retrieval through an embedding bottleneck. Failures occur when the generated text misses the query's true information need.

14

Conclusion

PAPER-DERIVED FACTS: HyDE shows that LLM generation can be used as an intermediate representation for zero-shot retrieval. It separates query understanding from evidence retrieval and became influential in query-transformation RAG pipelines.

15

References

PAPER-DERIVED FACTS: Primary source: Gao, L., Ma, X., Lin, J., & Callan, J. (2023), Precise Zero-Shot Dense Retrieval without Relevance Labels. ACL 2023. arXiv:2212.10496.

Continue reading

Related research

Browse all research