#math/calculus
A single-output derivative records how one output responds to the inputs. With several outputs, the local linear response needs every output-input sensitivity arranged in one table. For a map $F:\mathbb R^n\to\mathbb R^m$ with components $F_1,\ldots,F_m$, the `derivative matrix`, or `Jacobian`, is
$
J_F(x)=
\begin{bmatrix}
\frac{\partial F_1}{\partial x_1}&\cdots&\frac{\partial F_1}{\partial x_n}\\
\vdots&\ddots&\vdots\\
\frac{\partial F_m}{\partial x_1}&\cdots&\frac{\partial F_m}{\partial x_n}
\end{bmatrix}.
$
It has shape $m\times n$: one row per output and one column per input.^[The matrix shape corresponds to the matrices of linear maps defined in [[3C2 - Matrix of a Linear Map|LADR - Matrix of a Linear Map]].]
## Example: Jacobian of $F:\mathbb{R}^{2}\to \mathbb{R}^{3}$
Suppose $F:\mathbb{R}^{2}\to \mathbb{R}^{3}$ given by
$
F(x,y)=\langle u,v,w\rangle
=\langle x^2+y^2,\ x^2-y^2,\ xy\rangle.
$
Then, there are three outputs and two inputs, so there are 6 combinations of partial derivatives. Since $F:\mathbb{R}^{2}\to \mathbb{R}^{3}$, the derivative matrix must be $3\times2$:
$
J_F(x,y)=
\begin{array}{r@{\quad}c}
&
\begin{array}{c@{\qquad}c}
\scriptstyle \frac{\partial F}{\partial x}&
\scriptstyle \frac{\partial F}{\partial y}
\end{array}
\\[-1pt]
\begin{array}{r}
\vphantom{\dfrac{1}{1}}\scriptstyle \nabla u^{\mathsf T}=[u_x,u_y]\\
\vphantom{\dfrac{1}{1}}\scriptstyle \nabla v^{\mathsf T}=[v_x,v_y]\\
\vphantom{\dfrac{1}{1}}\scriptstyle \nabla w^{\mathsf T}=[w_x,w_y]
\end{array}
&
\left[
\begin{array}{c@{\qquad}c}
\vphantom{\dfrac{1}{1}}2x&2y\\
\vphantom{\dfrac{1}{1}}2x&-2y\\
\vphantom{\dfrac{1}{1}}y&x
\end{array}
\right]
\end{array}.
$
At $(-2,3)$,
$
J_F(-2,3)=
\begin{bmatrix}
-4&6\\
-4&-6\\
3&-2
\end{bmatrix}.
$
Each row tracks one output and each column tracks one input. Thus this single matrix records all six first-order sensitivities at the point.
At a point $p$, the best local linear model is
$
F(p+h)\approx F(p)+J_F(p)h.
$
Matrix multiplication is row-by-column: an $(m\times n)$ matrix times an $(n\times k)$ matrix produces an $(m\times k)$ matrix. The inner dimension must match. See [[3C3 - Matrix Multiplication|LADR - matrix multiplication]] for the underlying linear algebra.
> [!ML/AI]
> For a linear layer $h=Wx$, the Jacobian with respect to $x$ is $W$. If the next layer sends back an upstream gradient $g=\nabla_hL$, backprop computes $\nabla_xL=W^\top g$—a vector-Jacobian product—without storing a large network-wide Jacobian.