#math/calculus A single-output derivative records how one output responds to the inputs. With several outputs, the local linear response needs every output-input sensitivity arranged in one table. For a map $F:\mathbb R^n\to\mathbb R^m$ with components $F_1,\ldots,F_m$, the `derivative matrix`, or `Jacobian`, is $ J_F(x)= \begin{bmatrix} \frac{\partial F_1}{\partial x_1}&\cdots&\frac{\partial F_1}{\partial x_n}\\ \vdots&\ddots&\vdots\\ \frac{\partial F_m}{\partial x_1}&\cdots&\frac{\partial F_m}{\partial x_n} \end{bmatrix}. $ It has shape $m\times n$: one row per output and one column per input.^[The matrix shape corresponds to the matrices of linear maps defined in [[3C2 - Matrix of a Linear Map|LADR - Matrix of a Linear Map]].] ## Example: Jacobian of $F:\mathbb{R}^{2}\to \mathbb{R}^{3}$ Suppose $F:\mathbb{R}^{2}\to \mathbb{R}^{3}$ given by $ F(x,y)=\langle u,v,w\rangle =\langle x^2+y^2,\ x^2-y^2,\ xy\rangle. $ Then, there are three outputs and two inputs, so there are 6 combinations of partial derivatives. Since $F:\mathbb{R}^{2}\to \mathbb{R}^{3}$, the derivative matrix must be $3\times2$: $ J_F(x,y)= \begin{array}{r@{\quad}c} & \begin{array}{c@{\qquad}c} \scriptstyle \frac{\partial F}{\partial x}& \scriptstyle \frac{\partial F}{\partial y} \end{array} \\[-1pt] \begin{array}{r} \vphantom{\dfrac{1}{1}}\scriptstyle \nabla u^{\mathsf T}=[u_x,u_y]\\ \vphantom{\dfrac{1}{1}}\scriptstyle \nabla v^{\mathsf T}=[v_x,v_y]\\ \vphantom{\dfrac{1}{1}}\scriptstyle \nabla w^{\mathsf T}=[w_x,w_y] \end{array} & \left[ \begin{array}{c@{\qquad}c} \vphantom{\dfrac{1}{1}}2x&2y\\ \vphantom{\dfrac{1}{1}}2x&-2y\\ \vphantom{\dfrac{1}{1}}y&x \end{array} \right] \end{array}. $ At $(-2,3)$, $ J_F(-2,3)= \begin{bmatrix} -4&6\\ -4&-6\\ 3&-2 \end{bmatrix}. $ Each row tracks one output and each column tracks one input. Thus this single matrix records all six first-order sensitivities at the point. At a point $p$, the best local linear model is $ F(p+h)\approx F(p)+J_F(p)h. $ Matrix multiplication is row-by-column: an $(m\times n)$ matrix times an $(n\times k)$ matrix produces an $(m\times k)$ matrix. The inner dimension must match. See [[3C3 - Matrix Multiplication|LADR - matrix multiplication]] for the underlying linear algebra. > [!ML/AI] > For a linear layer $h=Wx$, the Jacobian with respect to $x$ is $W$. If the next layer sends back an upstream gradient $g=\nabla_hL$, backprop computes $\nabla_xL=W^\top g$—a vector-Jacobian product—without storing a large network-wide Jacobian.