Research

MachinoAI explainer / Retrieval-Augmented Generation

RAFT: Adapting Language Model to Domain Specific RAG

Fine-tunes an LLM for domain-specific RAG so it learns to identify and use relevant retrieved documents while ignoring distractors.

Tianjun Zhang, Shishir G. Patil, Naman Jain, Sheng Shen, Matei Zaharia, Ion Stoica, Joseph E. GonzalezMar 15, 2024arXiv 202412 min read
NEWRetrieval-Augmented GenerationAdvanced

01

Abstract

PAPER-DERIVED FACTS: RAFT fine-tunes a language model for domain-specific RAG by exposing it to relevant documents together with irrelevant retrieved distractors. The objective teaches the model to identify useful evidence and reason from it rather than treating every retrieved passage as equally trustworthy.

02

Introduction

PAPER-DERIVED FACTS: Domain-specific RAG systems often retrieve a mixture of useful and irrelevant documents. A general-purpose language model may not reliably distinguish them, motivating fine-tuning that explicitly teaches retrieval-aware reasoning.

03

Problem

PAPER-DERIVED FACTS: The model must answer questions using a retrieved corpus while ignoring distractors and, where appropriate, citing or reasoning from relevant passages. The training distribution should resemble the noisy evidence seen at inference time.

04

Background

PAPER-DERIVED FACTS: RAFT combines supervised fine-tuning with retrieval augmentation. It sits between closed-book fine-tuning, where knowledge is learned into weights, and pure open-book RAG, where the generator is not adapted to the retrieval environment.

05

Methodology

PAPER-DERIVED FACTS: Training examples contain a question, retrieved documents, and labels indicating which documents are useful. Relevant passages are paired with reasoning traces so the model learns to extract evidence and ignore distractors.

Figure notes

Visual evidence

Figure 1
RAFT retrieval-augmented fine-tuning workflow — RAFT retrieval-augmented fine-tuning workflowhttps://arxiv.org/pdf/2403.10131

06

Architecture

PAPER-DERIVED FACTS: At inference time, a standard retriever selects top-k domain documents and the fine-tuned model generates from them. The architectural novelty is mainly in the training procedure rather than a new retriever or decoder.

07

Dataset

PAPER-DERIVED FACTS: The study evaluates domain-specific RAG across datasets representing specialized knowledge. Training uses retrieval results containing both useful documents and sampled negative documents to mimic realistic retrieval noise.

08

Training

PAPER-DERIVED FACTS: RAFT fine-tunes an instruction-following language model on open-book question-answering examples. Distractor documents are intentionally included, and chain-of-thought-style reasoning can be used to teach the model how to identify supporting evidence.

09

Experiments

PAPER-DERIVED FACTS: The experiments compare RAFT against closed-book fine-tuning and standard RAG prompting across domain-specific QA tasks. The study examines both model adaptation and the effect of distractor-aware training.

10

Baselines

PAPER-DERIVED FACTS: Baselines include a base language model with RAG prompting and conventional supervised fine-tuning without retrieved documents. The comparisons test whether retrieval-aware fine-tuning adds value beyond simply giving the model more context.

11

Results

PAPER-DERIVED FACTS: RAFT improves domain-specific RAG performance on the reported tasks and shows that models trained with relevant documents plus distractors become better at using retrieval evidence under noisy conditions.

12

Ablation

PAPER-DERIVED FACTS: The paper studies the role of distractor documents, relevant passages, and reasoning supervision. The experiments indicate that learning to distinguish useful from irrelevant evidence is an important part of the gain.

13

Limitations

Error Analysis

PAPER-DERIVED FACTS: RAFT can still fail when the retriever omits the necessary evidence or when distractors contain plausible but misleading information. Fine-tuning cannot fully compensate for missing retrieval coverage.

14

Conclusion

PAPER-DERIVED FACTS: RAFT treats the generator itself as a component that can be optimized for noisy RAG conditions. The main lesson is that domain-specific RAG can benefit from training the model to reason over retrieved evidence rather than relying only on generic instruction following.

15

References

PAPER-DERIVED FACTS: Primary source: Zhang, T. et al. (2024), RAFT: Adapting Language Model to Domain Specific RAG. arXiv:2403.10131.

Continue reading

Related research

Browse all research