Research

MachinoAI explainer / Transformers

Attention Is All You Need

Introduces the Transformer: an attention-only sequence-transduction architecture that removes recurrence and enables highly parallel training.

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia PolosukhinDec 4, 2017NeurIPS 201725 min readGoogle Brain; Google Research; University of Toronto
FOUNDATIONALTransformersAdvancedscaled dot-product attentionmulti-head attentionpositional encoding

01

Abstract

PAPER-DERIVED FACTS: Attention Is All You Need introduced the Transformer, an encoder-decoder sequence-transduction architecture built entirely from attention mechanisms rather than recurrence or convolution. The paper targets a practical weakness of recurrent sequence models: computation is inherently sequential across token positions, which limits parallelization during training. The Transformer replaces that dependency chain with self-attention, allowing every position to interact directly within a layer. The paper reports 28.4 BLEU on WMT 2014 English-to-German for the Transformer big model and reports 41.8 BLEU in its abstract and Table 2 for English-to-French; the prose of Section 6.1 instead says 41.0 for the English-to-French big model, an internal inconsistency that this article preserves explicitly. The big model was trained for 3.5 days on eight NVIDIA P100 GPUs. The paper also evaluates English constituency parsing and reports 91.3 F1 in the WSJ-only setting and 92.7 F1 in the semi-supervised setting. ENGINEERING INTERPRETATION: The lasting contribution is not simply a better translation score. It is the separation of sequence modeling from sequential computation. Attention provides direct token-to-token communication while matrix multiplication provides hardware-friendly parallel computation.

02

Introduction

PAPER-DERIVED FACTS: Before the Transformer, strong sequence-to-sequence systems were dominated by RNN, LSTM, GRU, and convolutional architectures. Recurrent models process positions one after another, creating a serial bottleneck. Attention already existed, but was generally attached to recurrent encoders or decoders. The paper asks whether attention itself can become the sequence-transduction architecture. The answer is the Transformer. The important move is architectural: attention becomes the primary communication mechanism. An encoder position can read from all source positions, while a decoder uses masked self-attention and encoder-decoder attention. ENGINEERING INTERPRETATION: Training can process all known target positions in parallel under a causal mask. Generation is different: autoregressive decoding still produces one output token after another.

03

Problem

PAPER-DERIVED FACTS: The core problem is sequence transduction under three requirements: model long-range dependencies, train efficiently, and preserve autoregressive generation. Recurrent layers have O(n) sequential operations and O(n) maximum path length. Convolution can parallelize positions but needs multiple layers to connect distant positions. Self-attention instead has O(n^2 d) per-layer complexity, O(1) sequential operations during training, and O(1) maximum path length. The trade-off is explicit: dense attention is quadratic in sequence length. The paper recognizes restricted attention as a future direction for very long inputs. A second problem is positional information: removing recurrence and convolution removes an implicit notion of order, so positional encodings are added to embeddings. ENGINEERING INTERPRETATION: The formulation explains both the success and later research agenda. The original Transformer is highly parallel during training but does not solve long-context efficiency or autoregressive decoding latency.

04

Background

PAPER-DERIVED FACTS: The paper compares itself with Extended Neural GPU, ByteNet, ConvS2S, self-attention, memory networks, and attention-based sequence models. It evaluates computational complexity, minimum sequential operations, and maximum dependency path length. Self-attention gives direct connectivity between positions. The paper builds on additive and dot-product attention; dot-product attention is attractive because it maps efficiently to matrix multiplication. The 1/sqrt(d_k) scaling is motivated by the stated variance argument for dot products. The paper uses sinusoidal positional encoding and reports nearly identical results for learned positional embeddings in its ablation. EXTERNAL RESEARCH: BERT later used the Transformer encoder for bidirectional language representation learning; T5 explored a unified text-to-text transfer framework; Switch Transformer explored sparse expert scaling; FlashAttention attacked the memory-traffic bottleneck of exact attention. These works extend the Transformer rather than replacing its core abstraction.

05

Methodology

PAPER-DERIVED FACTS: The atomic operation is scaled dot-product attention: Attention(Q,K,V)=softmax(QK^T/sqrt(d_k))V. Queries determine what a position is looking for, keys provide comparable representations, and values contain the information aggregated. Multi-head attention projects Q, K and V into multiple lower-dimensional spaces, computes attention independently, concatenates the results, and applies a final projection. The base model uses eight heads with d_model=512 and d_k=d_v=64. Encoder layers contain multi-head self-attention and a position-wise feed-forward network. Decoder layers contain masked self-attention, encoder-decoder attention, and a feed-forward network. Residual connections and LayerNorm wrap each sub-layer using LayerNorm(x+Sublayer(x)). The feed-forward network is FFN(x)=max(0,xW1+b1)W2+b2, with d_ff=2048 in the base model. Embedding weights are shared with the pre-softmax linear transformation and scaled by sqrt(d_model).

Figure notes

Visual evidence

Figure 1
Figure 1 — The Transformer: model architecture — Original architecture diagram showing encoder and decoder stacks, embeddings, positional information, masked attention, encoder-decoder attention, feed-forward layers, residual Add & Norm operations, linear projection and softmax.Original Figure 1 from Vaswani et al. (2017), direct arXiv HTML image asset. Page 2
Figure 2
Figure 2A — Scaled Dot-Product Attention — Original attention primitive showing Q/K matrix multiplication, scaling, optional mask, softmax and multiplication by V.Original Figure 2 left panel from Vaswani et al. (2017), direct arXiv HTML image asset. Page 3
Figure 3
Figure 2B — Multi-Head Attention — Original multi-head diagram showing parallel projected attention operations, concatenation and final projection.Original Figure 2 right panel from Vaswani et al. (2017), direct arXiv HTML image asset. Page 3

Formula notes

Mathematical details

Scaled Dot-Product Attention

Equation
Attention(Q,K,V)=softmax(QKTdk)V\mathrm{Attention}(Q,K,V)=\mathrm{softmax}\left(\frac{QK^{T}}{\sqrt{d_k}}\right)V

Compute compatibility between queries and keys, normalize it, and use the resulting weights to mix the values.

K
key matrix
Q
query matrix
V
value matrix
dkd_k
key/query dimension

Multi-Head Attention

Equation
MultiHead(Q,K,V)=Concat(head1,…,headh)WO\mathrm{MultiHead}(Q,K,V)=\mathrm{Concat}(\mathrm{head}_1,\ldots,\mathrm{head}_h)W^O

Run several independently projected attention operations in parallel, concatenate their outputs, then project them back to model dimension.

h
number of heads
WOW^O
output projection
headihead_i
Attention(QW_i^Q,KW_i^K,VW_i^V)

Position-wise Feed-Forward Network

Equation
FFN(x)=max⁡(0,xW1+b1)W2+b2\mathrm{FFN}(x)=\max(0,xW_1+b_1)W_2+b_2

Apply the same two-layer nonlinear transformation independently to every token position within a layer.

W1W_1
first projection
W2W_2
second projection
b1b_1
first bias
b2b_2
second bias

Sinusoidal Positional Encoding

Equation
PE(pos,2i)=sin⁡(pos/100002i/dmodel)PE(pos,2i+1)=cos⁡(pos/100002i/dmodel)\begin{aligned}PE_{(pos,2i)} &= \sin(pos/10000^{2i/d_{model}}) \\[0.65em] PE_{(pos,2i+1)} &= \cos(pos/10000^{2i/d_{model}})\end{aligned}

Add deterministic position information to token embeddings so a recurrence-free architecture can represent sequence order.

i
dimension index
pos
token position
dmodeld_model
model dimension

Learning-Rate Schedule

Equation
lrate=dmodel−0.5min⁡(step_num−0.5,step_num⋅warmup_steps−1.5)lrate=d_{model}^{-0.5}\min(step\_num^{-0.5},step\_num\cdot warmup\_steps^{-1.5})

Warm up the learning rate linearly and then decay it proportional to the inverse square root of the training step.

dmodeld_model
model dimension
stepnumstep_num
optimization step
warmupstepswarmup_steps
4000

06

Architecture

PAPER-DERIVED FACTS: The canonical Transformer has six encoder layers and six decoder layers. Each encoder layer performs self-attention and then a position-wise feed-forward transformation. The decoder has masked self-attention, encoder-decoder attention, and a feed-forward network. Future target positions are masked before softmax, preserving autoregressive generation. Figure 1 is the main architectural map; Figure 2 decomposes attention into scaled dot-product and multi-head construction. ENGINEERING INTERPRETATION: The historical architecture should not be retrofitted with later components such as RoPE, RMSNorm, Pre-LN, KV caching, GQA, or FlashAttention. The original model is Post-LN, sinusoidal-positioned, encoder-decoder, and uses full multi-head attention. Keeping this boundary clear is essential when using the paper to learn modern LLMs.

07

Dataset

PAPER-DERIVED FACTS: The translation experiments use WMT 2014 English-German and English-French. English-German contains about 4.5 million sentence pairs and uses byte-pair encoding with a shared source-target vocabulary of about 37,000 tokens. English-French contains 36 million sentences and uses a 32,000-token word-piece vocabulary. Batches contain approximately 25,000 source and 25,000 target tokens. The parsing experiment uses the Wall Street Journal portion of Penn Treebank, about 40K training sentences in the WSJ-only setup, plus approximately 17M sentences in the semi-supervised setup. ENGINEERING INTERPRETATION: These datasets are small by modern LLM standards. The paper demonstrates a supervised sequence-transduction architecture, not large-scale unsupervised pretraining.

08

Training

PAPER-DERIVED FACTS: Training used one machine with eight NVIDIA P100 GPUs. Base steps took about 0.4 seconds for 100K steps, about 12 hours. Big-model steps took about 1.0 second for 300K steps, about 3.5 days. Adam used beta1=0.9, beta2=0.98 and epsilon=10^-9. The learning rate uses d_model^-0.5 times the minimum of step^-0.5 and step times warmup_steps^-1.5, with 4,000 warmup steps. Dropout was 0.1 for the base model and label smoothing was 0.1. Base models averaged the last five checkpoints; big models averaged the last twenty. Translation inference used beam size 4, length penalty alpha=0.6, and maximum output length input length plus 50. ENGINEERING INTERPRETATION: Reproducing the result requires the full recipe, not just the attention implementation.

09

Experiments

PAPER-DERIVED FACTS: The paper evaluates machine translation, architecture variations, and English constituency parsing. Translation uses BLEU and estimated training FLOPs. Architecture variations change head count, key/value dimensions, depth, feed-forward width, dropout, and positional encoding. Parsing uses a four-layer Transformer with d_model=1024 in WSJ-only and semi-supervised settings. The appendix contains qualitative attention visualizations of long-distance dependencies, anaphora, and different head behaviors. These visualizations are illustrative rather than systematic causal interpretability evidence.

10

Baselines

PAPER-DERIVED FACTS: Translation baselines include ByteNet, Deep-Attention plus PosUnk, GNMT+RL, ConvS2S, MoE, and ensemble variants. Table 2 reports English-German scores of 23.75, 24.6, 25.16, 26.03, 26.30, 26.36 for selected baselines, versus 27.3 Transformer base and 28.4 Transformer big. English-French baselines include 39.2, 39.92, 40.46, 40.56, 40.4, 41.16, 41.29, versus 38.1 base and 41.8 big in Table 2. ENGINEERING INTERPRETATION: These are literature baselines, not a fully controlled same-codebase benchmark; historical compute and tuning conditions differ.

11

Results

PAPER-DERIVED FACTS: Transformer big reaches 28.4 BLEU on WMT 2014 English-German and 41.8 in Table 2 for English-French. English-German training cost is reported as 2.3×10^19 FLOPs; the base model reports 3.3×10^18. The paper contains an important inconsistency: its current arXiv v7 abstract and Table 2 report 41.8 English-French BLEU, while Section 6.1 says 41.0. Google Research also displays 41.0 and the NeurIPS artifact contains the 41.0 wording. MachinoAI preserves both source values instead of silently choosing one. For the structured Table 2 record, 41.8 is retained because it is the current arXiv v7 table/abstract value. Parsing reports 91.3 F1 WSJ-only and 92.7 semi-supervised. ENGINEERING INTERPRETATION: The central result is improved translation quality with much more parallelizable training; autoregressive decoding remains sequential.

12

Ablation

PAPER-DERIVED FACTS: Table 3 uses English-German newstest2013 development data. Base is six layers, d_model=512, d_ff=2048, eight heads, d_k=d_v=64, dropout 0.1, label smoothing 0.1, 100K steps, PPL 4.92, BLEU 25.8, and 65M parameters. Head counts of 1, 4, 16 and 32 give BLEU 24.9, 25.5, 25.8 and 25.4. d_k values 16 and 32 give 25.1 and 25.4. Depths 2, 4 and 8 give 23.7, 25.3 and 25.5 in the listed configurations. Feed-forward widths 256, 1024 and 4096 give 24.5, 26.0 and 26.2. Learned positional embeddings give 4.92 PPL and 25.7 BLEU. The authors conclude that single-head and excessive-head settings can hurt quality, reduced key size hurts, larger models help, dropout helps avoid overfitting, and learned positions are nearly equivalent in this experiment.

13

Limitations

PAPER-DERIVED FACTS: Self-attention costs O(n^2 d) per layer. The paper proposes restricted attention for long inputs but leaves it for future work. Experiments focus on translation sentences and parsing rather than long-document reasoning. Generation remains autoregressive, and making generation less sequential is explicitly listed as future work. The sinusoidal extrapolation argument is a hypothesis, not a systematic long-context evaluation. Attention visualizations are qualitative and do not prove causal head specialization. Baselines are largely previously published systems, and FLOPs are estimated from hardware throughput. EXTERNAL RESEARCH: BERT, T5, Switch Transformer and FlashAttention address different limitations that remained after the original Transformer. FlashAttention improves memory traffic without changing mathematical exact attention; Switch Transformer scales parameters through sparse activation while introducing routing and communication costs.

14

Conclusion

PAPER-DERIVED FACTS: The paper introduced an attention-only sequence-transduction architecture using multi-head attention, feed-forward networks, residual connections, normalization and explicit positional information. It demonstrated strong translation results, faster parallel training and generalization to constituency parsing. ENGINEERING INTERPRETATION: Treat the paper as the architectural root of modern Transformers, not as a modern LLM specification. Decoder-only language modeling, BERT-style bidirectional pretraining, T5 text-to-text transfer, newer positional representations, normalization schemes, KV-sharing methods, optimized attention kernels and sparse experts are later developments. The most useful mental model is: Transformer = parallelizable token interaction + learned feature transformation + residual information flow + explicit position handling + causal masking where generation requires it.

15

References

Primary source: Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30, 5998–6008. The current arXiv record is version 7 of arXiv:1706.03762. Relevant paper-derived references include Bahdanau et al. on attention-based neural machine translation, He et al. on residual learning, Ba et al. on Layer Normalization, Kingma and Ba on Adam, Sennrich et al. on subword units, and Wu et al. on neural machine translation. EXTERNAL RESEARCH: Devlin et al. (2018), BERT; Raffel et al. (2020), T5; Fedus et al. (2022), Switch Transformers; and Dao et al. (2022), FlashAttention. These follow-up works show the research trajectory after the original paper without retrofitting later terminology into the 2017 architecture.

Primary links and external source pages used by this explainer.

Browse all research