#math/calculus
Separate coordinate slopes do not automatically guarantee that a surface behaves like a plane. `Differentiability` is the stronger claim that all small input changes fit one reliable local linear model. Near $p=(x_0,y_0)$,
$
f(x,y)\approx f(p)+f_x(p)(x-x_0)+f_y(p)(y-y_0).
$
Writing $\Delta x=x-x_0$ and $\Delta y=y-y_0$ and $\Delta f=f(x,y)-f(p)$, then
$
\Delta f\approx f_x(p)\Delta x+f_y(p)\Delta y.
$
More precisely, $f$ is differentiable at $p$ when the remainder is small relative to the input displacement:
$
f(p+h)-f(p)-Df(p)h=o(\|h\|),
$
where $p=(x_{0},y_{0})$ is the starting point; $h=(\Delta x,\Delta y)$ is the input displacement; $p+h$ is the point after making the change; $f(p+h)-f(p)$ is the actual change in $f$; $Df(p)h$^[In later notes, we will define whats given by $Df(p)^{\top}$ as the gradient.] is the change predicted by the linear approximation given by
$
Df(p)=[f_{x}(p)\quad f_{y}(p)];
$
and the predicted point after moving from $p\to h$ is given by
$
Df(p)h=
\begin{bmatrix}f_x(p)&f_y(p)\end{bmatrix}
\begin{bmatrix}\Delta x\\ \Delta y\end{bmatrix}
=f_x(p)\Delta x+f_y(p)\Delta y.
$
Continuous first partials near $p$ are a convenient sufficient condition.
![[Pasted image 20260928223725.png|400]]
## Example: Approximating ${(0.99e^{0.02})^8}$
Let
$
f(x,y)=(xe^y)^8=x^8e^{8y}.
$
The nearby point $(0.99,0.02)$ is close to the convenient base point $(1,0)$, where
$
f(1,0)=1,
\qquad
f_x(1,0)=8,
\qquad
f_y(1,0)=8.
$
Here $\Delta x=-0.01$ and $\Delta y=0.02$, so the linear approximation gives
$
f(0.99,0.02)
\approx 1+8(-0.01)+8(0.02)
=1.08.
$
A calculator gives approximately $1.082850933$. The small error shows the teaching of the note: near the base point, one linear model combines both input changes into a useful estimate.
The graph of this affine approximation is the [[15B2 - Tangent Planes and Normal Vectors|tangent plane]].
> [!ML/AI]
> Near parameters $\theta$, a loss obeys $L(\theta+\Delta)\approx L(\theta)+\nabla L(\theta)^\top\Delta$. Choosing $\Delta=-\eta\nabla L(\theta)$ predicts the decrease $-\eta\|\nabla L(\theta)\|^2$, which is the local justification for a gradient-descent step.