MachinoAI explainer / Retrieval-Augmented Generation
RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
Builds a hierarchical retrieval tree by recursively clustering and summarizing document chunks for multi-level retrieval.
01
Abstract
PAPER-DERIVED FACTS: RAPTOR constructs a tree of document representations by recursively embedding, clustering, and summarizing chunks. Retrieval can then select information at different levels of abstraction rather than only from short contiguous chunks.
02
Introduction
PAPER-DERIVED FACTS: Flat chunk retrieval is effective for local facts but can lose relationships and themes spanning a long document. RAPTOR introduces hierarchical summaries so retrieval can operate over both detailed and abstract representations.
03
Problem
PAPER-DERIVED FACTS: The system must answer questions whose evidence may be distributed across distant sections of a document. A flat retriever may retrieve isolated chunks without the broader context needed for multi-step reasoning.
04
Background
PAPER-DERIVED FACTS: RAPTOR builds on dense embeddings, clustering, abstractive summarization, and retrieval-augmented generation. Its main distinction is that summaries are recursively organized into a tree.
05
Methodology
PAPER-DERIVED FACTS: Leaf chunks are embedded and clustered; an LLM summarizes each cluster; the summaries are embedded and clustered again; this repeats until higher-level representations are formed. At inference, retrieval can traverse or collapse the tree.
06
Architecture
PAPER-DERIVED FACTS: The index is a multi-level tree containing original chunks at the leaves and progressively more abstract summaries above them. A query embedding is used to retrieve relevant nodes, which are then supplied to the language model.
07
Dataset
PAPER-DERIVED FACTS: The evaluation covers long-document question answering and benchmarks requiring multi-step reasoning, including QuALITY and other QA datasets reported in the paper.
08
Training
PAPER-DERIVED FACTS: RAPTOR does not require end-to-end retraining of the generator. Index construction uses embedding models and an LLM summarizer; the resulting tree can be paired with a downstream language model such as GPT-4.
09
Experiments
PAPER-DERIVED FACTS: Experiments compare tree-based retrieval with standard chunk retrieval and evaluate both retrieval configurations and final QA. The authors test recursive summaries on questions requiring information from different parts of documents.
10
Baselines
PAPER-DERIVED FACTS: Baselines include conventional flat retrieval systems that select contiguous chunks. The comparison tests whether hierarchical context provides a measurable advantage beyond simply retrieving more chunks.
11
Results
PAPER-DERIVED FACTS: RAPTOR improves performance on several long-document QA tasks and reports a 20 percentage-point absolute improvement on QuALITY when combined with GPT-4 over the previous best result cited by the authors.
12
Ablation
PAPER-DERIVED FACTS: The work compares tree traversal retrieval with collapsed-tree retrieval and studies the effect of hierarchical summaries. These experiments examine how abstraction level and retrieval strategy influence answer quality.
13
Limitations
Error Analysis
PAPER-DERIVED FACTS: Summaries can lose fine-grained details, while leaf chunks can lack global context. RAPTOR therefore faces a granularity trade-off: abstract nodes improve holistic understanding but may omit exact evidence.
14
Conclusion
PAPER-DERIVED FACTS: RAPTOR turns a flat document collection into a hierarchy of information at different abstraction levels. This enables RAG systems to retrieve both local facts and global summaries from long documents.
15
References
PAPER-DERIVED FACTS: Primary source: Sarthi, P. et al. (2024), RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval. ICLR 2024. arXiv:2401.18059.
Continue reading
Related research
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Defines the canonical RAG architecture that combines a dense retriever, an external document index, and a seq2seq generator trained end-to-end.
Retrieval-Augmented GenerationDense Passage Retrieval for Open-Domain Question Answering
Introduces a practical dual-encoder dense retriever that replaces sparse lexical matching with learned semantic passage representations.
Retrieval-Augmented GenerationREALM
Introduces end-to-end retrieval during language-model pre-training so factual knowledge can live in an external, replaceable corpus rather than only in model parameters.