Volume I
Mathematical Foundations
Linear algebra, calculus, probability, statistics, and optimization — the mathematical language of machine learning.
Linear Algebra
Vectors, Spans & Linear Independence
Volume I, Chapter 1 — Part I. A rigorous foundation for vector spaces, linear combinations, span, linear independence, inner products, and similarity. We prove the exchange lemma, develop the geometric language of machine learning representations, and derive the cosine-similarity framework underlying attention mechanisms.
Matrix Operations & Linear Transformations
Volume I, Chapter 1 — Part II. Matrices as linear transformations, composition, the four fundamental subspaces, rank-nullity theorem, determinants, and affine maps. Full derivations connecting matrix algebra to neural network layers and attention.
Eigenvalues & Eigenvectors
Volume I, Chapter 1 — Part III. Eigendecomposition, the characteristic polynomial, the spectral theorem for symmetric matrices, PCA as spectral optimization, and convergence of gradient descent in the eigenbasis.
Singular Value Decomposition (SVD)
Volume I, Chapter 1 — Part IV. Complete derivation of the SVD from the spectral theorem, Eckart-Young optimal low-rank approximation, connection to PCA, pseudoinverse, and theoretical foundations of LoRA and matrix completion.
Positive Definite Matrices
Volume I, Chapter 1 — Part V. Equivalent characterizations of positive definite and semidefinite matrices, Cholesky decomposition, convexity via the Hessian, covariance geometry, Schur complements, and regularization theory.
Calculus & Analysis
Probability Theory
Bayes' Theorem
Volume I, Chapter 3 — Part I. From Kolmogorov axioms to conditional probability, the product rule, Bayes' theorem, the law of total probability, sequential updating, and the Bayesian inference framework underlying probabilistic machine learning.
The Multivariate Gaussian Distribution
Volume I, Chapter 3 — Part II. Definition, characteristic function derivation, affine transformations, marginalization, conditioning via Schur complements, the information form, maximum entropy, and the role of Gaussians in Bayesian ML, Gaussian processes, and diffusion models.
KL Divergence
Volume I, Chapter 3 — Part IV. Kullback-Leibler divergence from cross-entropy and entropy, Gibbs' inequality proof, asymmetry and forward vs reverse KL, multivariate Gaussian closed form, and applications in VAEs, diffusion models, and RLHF.