MachinoAI explainer / Transformers
Attention Is All You Need
Introduces the Transformer: an attention-only sequence-transduction architecture that removes recurrence and enables highly parallel training.
01
Abstract
PAPER-DERIVED FACTS: Attention Is All You Need introduced the Transformer, an encoder-decoder sequence-transduction architecture built entirely from attention mechanisms rather than recurrence or convolution. The paper targets a practical weakness of recurrent sequence models: computation is inherently sequential across token positions, which limits parallelization during training. The Transformer replaces that dependency chain with self-attention, allowing every position to interact directly within a layer. The paper reports 28.4 BLEU on WMT 2014 English-to-German for the Transformer big model and reports 41.8 BLEU in its abstract and Table 2 for English-to-French; the prose of Section 6.1 instead says 41.0 for the English-to-French big model, an internal inconsistency that this article preserves explicitly. The big model was trained for 3.5 days on eight NVIDIA P100 GPUs. The paper also evaluates English constituency parsing and reports 91.3 F1 in the WSJ-only setting and 92.7 F1 in the semi-supervised setting. ENGINEERING INTERPRETATION: The lasting contribution is not simply a better translation score. It is the separation of sequence modeling from sequential computation. Attention provides direct token-to-token communication while matrix multiplication provides hardware-friendly parallel computation.
02
Introduction
PAPER-DERIVED FACTS: Before the Transformer, strong sequence-to-sequence systems were dominated by RNN, LSTM, GRU, and convolutional architectures. Recurrent models process positions one after another, creating a serial bottleneck. Attention already existed, but was generally attached to recurrent encoders or decoders. The paper asks whether attention itself can become the sequence-transduction architecture. The answer is the Transformer. The important move is architectural: attention becomes the primary communication mechanism. An encoder position can read from all source positions, while a decoder uses masked self-attention and encoder-decoder attention. ENGINEERING INTERPRETATION: Training can process all known target positions in parallel under a causal mask. Generation is different: autoregressive decoding still produces one output token after another.
03
Problem
PAPER-DERIVED FACTS: The core problem is sequence transduction under three requirements: model long-range dependencies, train efficiently, and preserve autoregressive generation. Recurrent layers have O(n) sequential operations and O(n) maximum path length. Convolution can parallelize positions but needs multiple layers to connect distant positions. Self-attention instead has O(n^2 d) per-layer complexity, O(1) sequential operations during training, and O(1) maximum path length. The trade-off is explicit: dense attention is quadratic in sequence length. The paper recognizes restricted attention as a future direction for very long inputs. A second problem is positional information: removing recurrence and convolution removes an implicit notion of order, so positional encodings are added to embeddings. ENGINEERING INTERPRETATION: The formulation explains both the success and later research agenda. The original Transformer is highly parallel during training but does not solve long-context efficiency or autoregressive decoding latency.
04
Background
PAPER-DERIVED FACTS: The paper compares itself with Extended Neural GPU, ByteNet, ConvS2S, self-attention, memory networks, and attention-based sequence models. It evaluates computational complexity, minimum sequential operations, and maximum dependency path length. Self-attention gives direct connectivity between positions. The paper builds on additive and dot-product attention; dot-product attention is attractive because it maps efficiently to matrix multiplication. The 1/sqrt(d_k) scaling is motivated by the stated variance argument for dot products. The paper uses sinusoidal positional encoding and reports nearly identical results for learned positional embeddings in its ablation. EXTERNAL RESEARCH: BERT later used the Transformer encoder for bidirectional language representation learning; T5 explored a unified text-to-text transfer framework; Switch Transformer explored sparse expert scaling; FlashAttention attacked the memory-traffic bottleneck of exact attention. These works extend the Transformer rather than replacing its core abstraction.
05
Methodology
PAPER-DERIVED FACTS: The atomic operation is scaled dot-product attention: Attention(Q,K,V)=softmax(QK^T/sqrt(d_k))V. Queries determine what a position is looking for, keys provide comparable representations, and values contain the information aggregated. Multi-head attention projects Q, K and V into multiple lower-dimensional spaces, computes attention independently, concatenates the results, and applies a final projection. The base model uses eight heads with d_model=512 and d_k=d_v=64. Encoder layers contain multi-head self-attention and a position-wise feed-forward network. Decoder layers contain masked self-attention, encoder-decoder attention, and a feed-forward network. Residual connections and LayerNorm wrap each sub-layer using LayerNorm(x+Sublayer(x)). The feed-forward network is FFN(x)=max(0,xW1+b1)W2+b2, with d_ff=2048 in the base model. Embedding weights are shared with the pre-softmax linear transformation and scaled by sqrt(d_model).
Figure notes
Visual evidence
Formula notes
Mathematical details
Scaled Dot-Product Attention
Compute compatibility between queries and keys, normalize it, and use the resulting weights to mix the values.
- K
- key matrix
- Q
- query matrix
- V
- value matrix
- key/query dimension
Multi-Head Attention
Run several independently projected attention operations in parallel, concatenate their outputs, then project them back to model dimension.
- h
- number of heads
- output projection
- Attention(QW_i^Q,KW_i^K,VW_i^V)
Position-wise Feed-Forward Network
Apply the same two-layer nonlinear transformation independently to every token position within a layer.
- first projection
- second projection
- first bias
- second bias
Sinusoidal Positional Encoding
Add deterministic position information to token embeddings so a recurrence-free architecture can represent sequence order.
- i
- dimension index
- pos
- token position
- model dimension
Learning-Rate Schedule
Warm up the learning rate linearly and then decay it proportional to the inverse square root of the training step.
- model dimension
- optimization step
- 4000
06
Architecture
PAPER-DERIVED FACTS: The canonical Transformer has six encoder layers and six decoder layers. Each encoder layer performs self-attention and then a position-wise feed-forward transformation. The decoder has masked self-attention, encoder-decoder attention, and a feed-forward network. Future target positions are masked before softmax, preserving autoregressive generation. Figure 1 is the main architectural map; Figure 2 decomposes attention into scaled dot-product and multi-head construction. ENGINEERING INTERPRETATION: The historical architecture should not be retrofitted with later components such as RoPE, RMSNorm, Pre-LN, KV caching, GQA, or FlashAttention. The original model is Post-LN, sinusoidal-positioned, encoder-decoder, and uses full multi-head attention. Keeping this boundary clear is essential when using the paper to learn modern LLMs.
07
Dataset
PAPER-DERIVED FACTS: The translation experiments use WMT 2014 English-German and English-French. English-German contains about 4.5 million sentence pairs and uses byte-pair encoding with a shared source-target vocabulary of about 37,000 tokens. English-French contains 36 million sentences and uses a 32,000-token word-piece vocabulary. Batches contain approximately 25,000 source and 25,000 target tokens. The parsing experiment uses the Wall Street Journal portion of Penn Treebank, about 40K training sentences in the WSJ-only setup, plus approximately 17M sentences in the semi-supervised setup. ENGINEERING INTERPRETATION: These datasets are small by modern LLM standards. The paper demonstrates a supervised sequence-transduction architecture, not large-scale unsupervised pretraining.
08
Training
PAPER-DERIVED FACTS: Training used one machine with eight NVIDIA P100 GPUs. Base steps took about 0.4 seconds for 100K steps, about 12 hours. Big-model steps took about 1.0 second for 300K steps, about 3.5 days. Adam used beta1=0.9, beta2=0.98 and epsilon=10^-9. The learning rate uses d_model^-0.5 times the minimum of step^-0.5 and step times warmup_steps^-1.5, with 4,000 warmup steps. Dropout was 0.1 for the base model and label smoothing was 0.1. Base models averaged the last five checkpoints; big models averaged the last twenty. Translation inference used beam size 4, length penalty alpha=0.6, and maximum output length input length plus 50. ENGINEERING INTERPRETATION: Reproducing the result requires the full recipe, not just the attention implementation.
09
Experiments
PAPER-DERIVED FACTS: The paper evaluates machine translation, architecture variations, and English constituency parsing. Translation uses BLEU and estimated training FLOPs. Architecture variations change head count, key/value dimensions, depth, feed-forward width, dropout, and positional encoding. Parsing uses a four-layer Transformer with d_model=1024 in WSJ-only and semi-supervised settings. The appendix contains qualitative attention visualizations of long-distance dependencies, anaphora, and different head behaviors. These visualizations are illustrative rather than systematic causal interpretability evidence.
10
Baselines
PAPER-DERIVED FACTS: Translation baselines include ByteNet, Deep-Attention plus PosUnk, GNMT+RL, ConvS2S, MoE, and ensemble variants. Table 2 reports English-German scores of 23.75, 24.6, 25.16, 26.03, 26.30, 26.36 for selected baselines, versus 27.3 Transformer base and 28.4 Transformer big. English-French baselines include 39.2, 39.92, 40.46, 40.56, 40.4, 41.16, 41.29, versus 38.1 base and 41.8 big in Table 2. ENGINEERING INTERPRETATION: These are literature baselines, not a fully controlled same-codebase benchmark; historical compute and tuning conditions differ.
11
Results
PAPER-DERIVED FACTS: Transformer big reaches 28.4 BLEU on WMT 2014 English-German and 41.8 in Table 2 for English-French. English-German training cost is reported as 2.3×10^19 FLOPs; the base model reports 3.3×10^18. The paper contains an important inconsistency: its current arXiv v7 abstract and Table 2 report 41.8 English-French BLEU, while Section 6.1 says 41.0. Google Research also displays 41.0 and the NeurIPS artifact contains the 41.0 wording. MachinoAI preserves both source values instead of silently choosing one. For the structured Table 2 record, 41.8 is retained because it is the current arXiv v7 table/abstract value. Parsing reports 91.3 F1 WSJ-only and 92.7 semi-supervised. ENGINEERING INTERPRETATION: The central result is improved translation quality with much more parallelizable training; autoregressive decoding remains sequential.
12
Ablation
PAPER-DERIVED FACTS: Table 3 uses English-German newstest2013 development data. Base is six layers, d_model=512, d_ff=2048, eight heads, d_k=d_v=64, dropout 0.1, label smoothing 0.1, 100K steps, PPL 4.92, BLEU 25.8, and 65M parameters. Head counts of 1, 4, 16 and 32 give BLEU 24.9, 25.5, 25.8 and 25.4. d_k values 16 and 32 give 25.1 and 25.4. Depths 2, 4 and 8 give 23.7, 25.3 and 25.5 in the listed configurations. Feed-forward widths 256, 1024 and 4096 give 24.5, 26.0 and 26.2. Learned positional embeddings give 4.92 PPL and 25.7 BLEU. The authors conclude that single-head and excessive-head settings can hurt quality, reduced key size hurts, larger models help, dropout helps avoid overfitting, and learned positions are nearly equivalent in this experiment.
13
Limitations
PAPER-DERIVED FACTS: Self-attention costs O(n^2 d) per layer. The paper proposes restricted attention for long inputs but leaves it for future work. Experiments focus on translation sentences and parsing rather than long-document reasoning. Generation remains autoregressive, and making generation less sequential is explicitly listed as future work. The sinusoidal extrapolation argument is a hypothesis, not a systematic long-context evaluation. Attention visualizations are qualitative and do not prove causal head specialization. Baselines are largely previously published systems, and FLOPs are estimated from hardware throughput. EXTERNAL RESEARCH: BERT, T5, Switch Transformer and FlashAttention address different limitations that remained after the original Transformer. FlashAttention improves memory traffic without changing mathematical exact attention; Switch Transformer scales parameters through sparse activation while introducing routing and communication costs.
14
Conclusion
PAPER-DERIVED FACTS: The paper introduced an attention-only sequence-transduction architecture using multi-head attention, feed-forward networks, residual connections, normalization and explicit positional information. It demonstrated strong translation results, faster parallel training and generalization to constituency parsing. ENGINEERING INTERPRETATION: Treat the paper as the architectural root of modern Transformers, not as a modern LLM specification. Decoder-only language modeling, BERT-style bidirectional pretraining, T5 text-to-text transfer, newer positional representations, normalization schemes, KV-sharing methods, optimized attention kernels and sparse experts are later developments. The most useful mental model is: Transformer = parallelizable token interaction + learned feature transformation + residual information flow + explicit position handling + causal masking where generation requires it.
15
References
Primary source: Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30, 5998–6008. The current arXiv record is version 7 of arXiv:1706.03762. Relevant paper-derived references include Bahdanau et al. on attention-based neural machine translation, He et al. on residual learning, Ba et al. on Layer Normalization, Kingma and Ba on Adam, Sennrich et al. on subword units, and Wu et al. on neural machine translation. EXTERNAL RESEARCH: Devlin et al. (2018), BERT; Raffel et al. (2020), T5; Fedus et al. (2022), Switch Transformers; and Dao et al. (2022), FlashAttention. These follow-up works show the research trajectory after the original paper without retrofitting later terminology into the 2017 architecture.
Primary links and external source pages used by this explainer.
Attention Is All You Need
Vaswani et al. (2017)
Primary source.
Attention Is All You Need
Vaswani et al. (2017), NeurIPS 2017
Conference artifact.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin et al. (2018)
Transformer encoder and bidirectional pretraining.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel et al. (2020)
Text-to-text Transformer transfer learning.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Fedus et al. (2022)
Sparse expert scaling.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Dao et al. (2022)
IO-aware exact attention.
arXiv full paperOriginal PDFarXiv HTML renderingNeurIPS 2017 paperGoogle Research publicationBERTT5Switch TransformerFlashAttentionTensor2Tensor code