#math/calculus
[[15D1 - Derivative Matrices]] constructs the derivative matrix $J_F(p)$ of one map, so $\Delta F\approx J_F(p)\Delta x$. Here we differentiate a composition $f\circ g$ by composing the derivative matrices of its stages. Let $u\in\mathbb R^k$ be the initial input, $x=g(u)\in\mathbb R^n$ the intermediate value, and $y=f(x)\in\mathbb R^m$ the final output. Their first-order changes pass through
$
\Delta u\xmapsto{\,J_g(u)\,}\Delta x
\xmapsto{\,J_f(g(u))\,}\Delta y.
$
So, given two function $g:\mathbb R^k\to\mathbb R^n$ and $f:\mathbb R^n\to\mathbb R^m$. We define the Jacobian of their composition, also known as the `chain rule for jacobians`, or the `matrix chain rule` by
$
\boxed{J_{f\circ g}(u)=J_f(g(u))J_g(u)}.
$
The order of the matrix factors matters: the rightmost derivative acts first.
Entrywise, the value of $J_{f\circ g}$ at row $i$ column $j$ is given by
$
\frac{\partial(f\circ g)_i}{\partial u_j}
=\sum_{\ell=1}^{n}
\frac{\partial f_i}{\partial x_\ell}
\frac{\partial g_\ell}{\partial u_j}.
$
This is the scalar chain rule with every intermediate variable included.
## Example: Dummy Jacobian Composition
Let
$
g:\mathbb R^2\to\mathbb R^2,
\qquad
g(u,v)=\bigl(x(u,v),y(u,v)\bigr),
$
and
$
f:\mathbb R^2\to\mathbb R^2,
\qquad
f(x,y)=\bigl(r(x,y),s(x,y)\bigr).
$
Write $(f\circ g)(u,v)=\bigl(R(u,v),S(u,v)\bigr)$, where $R=r\circ g$ and $S=s\circ g$. Their Jacobians compose as
$
\underbrace{
\begin{bmatrix}
\dfrac{\partial R}{\partial u}&\dfrac{\partial R}{\partial v}\\[6pt]
\dfrac{\partial S}{\partial u}&\dfrac{\partial S}{\partial v}
\end{bmatrix}}_{J_{f\circ g}(u,v)}
=
\underbrace{
\begin{bmatrix}
\dfrac{\partial r}{\partial x}&\dfrac{\partial r}{\partial y}\\[6pt]
\dfrac{\partial s}{\partial x}&\dfrac{\partial s}{\partial y}
\end{bmatrix}}_{J_f(g(u,v))}
\underbrace{
\begin{bmatrix}
\dfrac{\partial x}{\partial u}&\dfrac{\partial x}{\partial v}\\[6pt]
\dfrac{\partial y}{\partial u}&\dfrac{\partial y}{\partial v}
\end{bmatrix}}_{J_g(u,v)}.
$
Multiplying row by column gives
$
J_{f\circ g}(u,v)=
\begin{bmatrix}
\dfrac{\partial r}{\partial x}\dfrac{\partial x}{\partial u}
+\dfrac{\partial r}{\partial y}\dfrac{\partial y}{\partial u}
&
\dfrac{\partial r}{\partial x}\dfrac{\partial x}{\partial v}
+\dfrac{\partial r}{\partial y}\dfrac{\partial y}{\partial v}
\\[10pt]
\dfrac{\partial s}{\partial x}\dfrac{\partial x}{\partial u}
+\dfrac{\partial s}{\partial y}\dfrac{\partial y}{\partial u}
&
\dfrac{\partial s}{\partial x}\dfrac{\partial x}{\partial v}
+\dfrac{\partial s}{\partial y}\dfrac{\partial y}{\partial v}
\end{bmatrix}.
$
The derivatives of $r$ and $s$ are evaluated at $(x,y)=g(u,v)$.
## Example: Linear Change of Basis
Suppose
$
(t,s)=(f(x,y),g(x,y)),
\qquad
x=u-2v,
\qquad
y=u+3v.
$
The two maps are
$
\underbrace{(u,v)\longmapsto(x,y)=(u-2v,u+3v)}_{\text{inner map: linear change of coordinates}},
\qquad
\underbrace{(x,y)\longmapsto(t,s)=(f(x,y),g(x,y))}_{\text{outer map}}.
$
The inner coordinate change has Jacobian
$
\frac{\partial(x,y)}{\partial(u,v)}
=
\begin{bmatrix}
1&-2\\
1&3
\end{bmatrix}.
$
The matrix chain rule then gives
$
\frac{\partial(t,s)}{\partial(u,v)}
=
\frac{\partial(t,s)}{\partial(x,y)}
\frac{\partial(x,y)}{\partial(u,v)}
=
\begin{bmatrix}
t_x&t_y\\
s_x&s_y
\end{bmatrix}
\begin{bmatrix}
1&-2\\
1&3
\end{bmatrix},
$
so
$
\frac{\partial(t,s)}{\partial(u,v)}
=
\begin{bmatrix}
t_x+t_y&-2t_x+3t_y\\
s_x+s_y&-2s_x+3s_y
\end{bmatrix},
$
with the outer derivatives evaluated at $(x,y)=(u-2v,u+3v)$. The product order follows the computation: perturb $(u,v)$, convert it to a perturbation in $(x,y)$, and only then measure the response in $(t,s)$.
For polar coordinates $g(r,\theta)=(r\cos\theta,r\sin\theta)$,
$
J_g=
\begin{bmatrix}
\cos\theta&-r\sin\theta\\
\sin\theta&r\cos\theta
\end{bmatrix},
\qquad
\det J_g=r.
$
The determinant measures local area scaling: a small $dr\,d\theta$ rectangle becomes area approximately $r\,dr\,d\theta$.
> [!ML/AI]
> Consider a two-layer neural network $F:\mathbb R^n\to\mathbb R^m$ defined by
> $
> z=W_1x,
> \qquad
> h=\operatorname{ReLU}(z),
> \qquad
> F(x)=W_2h.
> $
> Here $x$ is the input, $W_1$ and $W_2$ are weight matrices, $z$ contains the hidden layer's pre-activation values, and $h$ contains its activations. If $D_{\operatorname{ReLU}}(z)$ is the diagonal matrix of ReLU derivatives, then the matrix chain rule gives
> $
> J_F(x)=W_2D_{\operatorname{ReLU}}(z)W_1.
> $
> For a scalar loss $L:\mathbb R^m\to\mathbb R$, let $\delta=\nabla_F L$ be the gradient arriving from the loss. Backpropagation applies the transposed Jacobians in reverse order:
> $
> \nabla_xL
> =J_F(x)^\top\delta
> =W_1^\top D_{\operatorname{ReLU}}(z)W_2^\top\delta.
> $
> Thus the gradient passes backward through $W_2$, is masked by the active ReLUs, and then passes through $W_1$. This is the Jacobian chain rule implemented without forming the full Jacobian matrix.