Rotary Position Embeddings (RoPE)
Volume III, Chapter 11 — Part II. Complete derivation of RoPE: rotation matrices, relative position in attention scores, frequency spectrum, length extrapolation, and NTK-aware scaling.
Table of Contents
- Learning Objectives
- Prerequisites
- Notation
- Core Intuition
- Rotation Matrices and 2D Blocks
- RoPE Definition
- Relative Position Property
- Frequency Spectrum and Distance Decay
- Implementation Structure
- Length Extrapolation and NTK Scaling
- Comparison with Other Position Methods
- Worked Examples
- Connection to the Broader Curriculum
- Common Pitfalls and Misconceptions
- Research Perspective
- Summary of Takeaways
- Exercises
Learning Objectives
After reading this chapter, you should be able to:
- Define RoPE as block-diagonal rotation matrices applied to Q and K.
- Prove .
- Explain the geometric frequency spectrum and long-range decay.
- Describe NTK-aware scaling for context extension.
- Explain RoPE compatibility with KV Cache.
Prerequisites
- Positional Encoding — sinusoidal PE, rotation derivation
- Self-Attention — Q/K/V
- Eigenvalues & Eigenvectors — rotation as orthogonal transform
Notation
- — Absolute positions in sequence
- — Head dimension (assumed even)
- — Rotary frequency at dimension pair
- — Rotation matrix applied at position
- — Query and key vectors with RoPE applied
Core Intuition
Positional Encoding showed that encoding position in Q/K dot products enables relative position awareness. RoPE (Su et al., 2021) applies this directly: rotate Q and K by position-dependent angles before computing attention.
The attention score depends only on relative position , not absolute indices. RoPE requires zero extra parameters, works with KV Cache, and extrapolates to longer sequences with appropriate scaling — making it the standard in LLaMA, Mistral, and GPT-NeoX.
Series context. Volume III, Chapter 11, Part II. Extends Positional Encoding.
Rotary Position Embeddings (RoPE)
Rotation Matrices and 2D Blocks
Definition 1 (2D Rotation).
Proposition 1. .
Proof. Rotation composition adds angles; transpose inverts rotation.
RoPE Definition
Definition 2 (RoPE Frequencies).
Typical .
Definition 3 (RoPE Transform). For position , apply block-diagonal rotation:
Attention uses rotated Q, K; V is unchanged.
Relative Position Property
Theorem 1 (RoPE Relative Attention).
Proof.
by Proposition 1.
Important equation. Attention depends on only — native relative position encoding.
Frequency Spectrum and Distance Decay
Proposition 2. High-frequency components ( large) encode fine local position; low-frequency components encode coarse global position.
Proposition 3. For large , high-frequency terms oscillate rapidly, reducing correlation — implicit distance decay.
Analogous to sinusoidal PE in Positional Encoding.
Implementation Structure
RoPE acts on consecutive dimension pairs :
For cached K: store or apply rotation at cache-write time with position index.
Length Extrapolation and NTK Scaling
Problem. RoPE trained on length may degrade at .
Definition 4 (NTK-Aware Scaling). Rescale base frequency at inference:
Interpretation. Interpolates between frequencies to maintain attention coherence at extended context.
Related: Position interpolation (PI), YaRN — active research for million-token contexts.
Comparison with Other Position Methods
See Positional Encoding comparison table. RoPE advantages:
- Zero parameters
- Native relative encoding
- KV-cache compatible
- Better extrapolation than learned absolute PE
Worked Examples
Example 1: Relative Invariance
Attention between positions equals — both distance 3.
Example 2: Rotation Angle
At , may exceed — periodicity wraps position information.
Connection to the Broader Curriculum
- Positional Encoding — predecessor
- Self-Attention — application
- KV Cache — inference
- Flash Attention — efficient attention with RoPE
Common Pitfalls and Misconceptions
Pitfall 1: Applying RoPE to V (incorrect — only Q and K).
Pitfall 2: Wrong position index during cached decode.
Pitfall 3: Assuming unlimited extrapolation without scaling.
Research Perspective
RoPE (Su et al., 2021). NTK scaling (bloc97, 2023). YaRN, LongRoPE. Alternatives: ALiBi, CoPE.
Summary of Takeaways
- RoPE —
- Relative —
- Frequencies —
- Extrapolation — NTK scaling
Next: Scaling Laws
Exercises
Exercise 1. Prove Proposition 1.
Exercise 2. Prove Theorem 1.
Exercise 3. Derive (7) for single pair.
Exercise 4. Why is V not rotated?
Exercise 5. KV cache: when to apply ?
Exercise 6. Compare RoPE to sinusoidal PE addition.
Exercise 7. Effect of on frequency spectrum.
Exercise 8. Derive NTK scaling motivation heuristically.