#math/calculus [[15D1 - Derivative Matrices]] constructs the derivative matrix $J_F(p)$ of one map, so $\Delta F\approx J_F(p)\Delta x$. Here we differentiate a composition $f\circ g$ by composing the derivative matrices of its stages. Let $u\in\mathbb R^k$ be the initial input, $x=g(u)\in\mathbb R^n$ the intermediate value, and $y=f(x)\in\mathbb R^m$ the final output. Their first-order changes pass through $ \Delta u\xmapsto{\,J_g(u)\,}\Delta x \xmapsto{\,J_f(g(u))\,}\Delta y. $ So, given two function $g:\mathbb R^k\to\mathbb R^n$ and $f:\mathbb R^n\to\mathbb R^m$. We define the Jacobian of their composition, also known as the `chain rule for jacobians`, or the `matrix chain rule` by $ \boxed{J_{f\circ g}(u)=J_f(g(u))J_g(u)}. $ The order of the matrix factors matters: the rightmost derivative acts first. Entrywise, the value of $J_{f\circ g}$ at row $i$ column $j$ is given by $ \frac{\partial(f\circ g)_i}{\partial u_j} =\sum_{\ell=1}^{n} \frac{\partial f_i}{\partial x_\ell} \frac{\partial g_\ell}{\partial u_j}. $ This is the scalar chain rule with every intermediate variable included. ## Example: Dummy Jacobian Composition Let $ g:\mathbb R^2\to\mathbb R^2, \qquad g(u,v)=\bigl(x(u,v),y(u,v)\bigr), $ and $ f:\mathbb R^2\to\mathbb R^2, \qquad f(x,y)=\bigl(r(x,y),s(x,y)\bigr). $ Write $(f\circ g)(u,v)=\bigl(R(u,v),S(u,v)\bigr)$, where $R=r\circ g$ and $S=s\circ g$. Their Jacobians compose as $ \underbrace{ \begin{bmatrix} \dfrac{\partial R}{\partial u}&\dfrac{\partial R}{\partial v}\\[6pt] \dfrac{\partial S}{\partial u}&\dfrac{\partial S}{\partial v} \end{bmatrix}}_{J_{f\circ g}(u,v)} = \underbrace{ \begin{bmatrix} \dfrac{\partial r}{\partial x}&\dfrac{\partial r}{\partial y}\\[6pt] \dfrac{\partial s}{\partial x}&\dfrac{\partial s}{\partial y} \end{bmatrix}}_{J_f(g(u,v))} \underbrace{ \begin{bmatrix} \dfrac{\partial x}{\partial u}&\dfrac{\partial x}{\partial v}\\[6pt] \dfrac{\partial y}{\partial u}&\dfrac{\partial y}{\partial v} \end{bmatrix}}_{J_g(u,v)}. $ Multiplying row by column gives $ J_{f\circ g}(u,v)= \begin{bmatrix} \dfrac{\partial r}{\partial x}\dfrac{\partial x}{\partial u} +\dfrac{\partial r}{\partial y}\dfrac{\partial y}{\partial u} & \dfrac{\partial r}{\partial x}\dfrac{\partial x}{\partial v} +\dfrac{\partial r}{\partial y}\dfrac{\partial y}{\partial v} \\[10pt] \dfrac{\partial s}{\partial x}\dfrac{\partial x}{\partial u} +\dfrac{\partial s}{\partial y}\dfrac{\partial y}{\partial u} & \dfrac{\partial s}{\partial x}\dfrac{\partial x}{\partial v} +\dfrac{\partial s}{\partial y}\dfrac{\partial y}{\partial v} \end{bmatrix}. $ The derivatives of $r$ and $s$ are evaluated at $(x,y)=g(u,v)$. ## Example: Linear Change of Basis Suppose $ (t,s)=(f(x,y),g(x,y)), \qquad x=u-2v, \qquad y=u+3v. $ The two maps are $ \underbrace{(u,v)\longmapsto(x,y)=(u-2v,u+3v)}_{\text{inner map: linear change of coordinates}}, \qquad \underbrace{(x,y)\longmapsto(t,s)=(f(x,y),g(x,y))}_{\text{outer map}}. $ The inner coordinate change has Jacobian $ \frac{\partial(x,y)}{\partial(u,v)} = \begin{bmatrix} 1&-2\\ 1&3 \end{bmatrix}. $ The matrix chain rule then gives $ \frac{\partial(t,s)}{\partial(u,v)} = \frac{\partial(t,s)}{\partial(x,y)} \frac{\partial(x,y)}{\partial(u,v)} = \begin{bmatrix} t_x&t_y\\ s_x&s_y \end{bmatrix} \begin{bmatrix} 1&-2\\ 1&3 \end{bmatrix}, $ so $ \frac{\partial(t,s)}{\partial(u,v)} = \begin{bmatrix} t_x+t_y&-2t_x+3t_y\\ s_x+s_y&-2s_x+3s_y \end{bmatrix}, $ with the outer derivatives evaluated at $(x,y)=(u-2v,u+3v)$. The product order follows the computation: perturb $(u,v)$, convert it to a perturbation in $(x,y)$, and only then measure the response in $(t,s)$. For polar coordinates $g(r,\theta)=(r\cos\theta,r\sin\theta)$, $ J_g= \begin{bmatrix} \cos\theta&-r\sin\theta\\ \sin\theta&r\cos\theta \end{bmatrix}, \qquad \det J_g=r. $ The determinant measures local area scaling: a small $dr\,d\theta$ rectangle becomes area approximately $r\,dr\,d\theta$. > [!ML/AI] > Consider a two-layer neural network $F:\mathbb R^n\to\mathbb R^m$ defined by > $ > z=W_1x, > \qquad > h=\operatorname{ReLU}(z), > \qquad > F(x)=W_2h. > $ > Here $x$ is the input, $W_1$ and $W_2$ are weight matrices, $z$ contains the hidden layer's pre-activation values, and $h$ contains its activations. If $D_{\operatorname{ReLU}}(z)$ is the diagonal matrix of ReLU derivatives, then the matrix chain rule gives > $ > J_F(x)=W_2D_{\operatorname{ReLU}}(z)W_1. > $ > For a scalar loss $L:\mathbb R^m\to\mathbb R$, let $\delta=\nabla_F L$ be the gradient arriving from the loss. Backpropagation applies the transposed Jacobians in reverse order: > $ > \nabla_xL > =J_F(x)^\top\delta > =W_1^\top D_{\operatorname{ReLU}}(z)W_2^\top\delta. > $ > Thus the gradient passes backward through $W_2$, is masked by the active ReLUs, and then passes through $W_1$. This is the Jacobian chain rule implemented without forming the full Jacobian matrix.