#math/calculus Separate coordinate slopes do not automatically guarantee that a surface behaves like a plane. `Differentiability` is the stronger claim that all small input changes fit one reliable local linear model. Near $p=(x_0,y_0)$, $ f(x,y)\approx f(p)+f_x(p)(x-x_0)+f_y(p)(y-y_0). $ Writing $\Delta x=x-x_0$ and $\Delta y=y-y_0$ and $\Delta f=f(x,y)-f(p)$, then $ \Delta f\approx f_x(p)\Delta x+f_y(p)\Delta y. $ More precisely, $f$ is differentiable at $p$ when the remainder is small relative to the input displacement: $ f(p+h)-f(p)-Df(p)h=o(\|h\|), $ where $p=(x_{0},y_{0})$ is the starting point; $h=(\Delta x,\Delta y)$ is the input displacement; $p+h$ is the point after making the change; $f(p+h)-f(p)$ is the actual change in $f$; $Df(p)h$^[In later notes, we will define whats given by $Df(p)^{\top}$ as the gradient.] is the change predicted by the linear approximation given by $ Df(p)=[f_{x}(p)\quad f_{y}(p)]; $ and the predicted point after moving from $p\to h$ is given by $ Df(p)h= \begin{bmatrix}f_x(p)&f_y(p)\end{bmatrix} \begin{bmatrix}\Delta x\\ \Delta y\end{bmatrix} =f_x(p)\Delta x+f_y(p)\Delta y. $ Continuous first partials near $p$ are a convenient sufficient condition. ![[Pasted image 20260928223725.png|400]] ## Example: Approximating ${(0.99e^{0.02})^8}$ Let $ f(x,y)=(xe^y)^8=x^8e^{8y}. $ The nearby point $(0.99,0.02)$ is close to the convenient base point $(1,0)$, where $ f(1,0)=1, \qquad f_x(1,0)=8, \qquad f_y(1,0)=8. $ Here $\Delta x=-0.01$ and $\Delta y=0.02$, so the linear approximation gives $ f(0.99,0.02) \approx 1+8(-0.01)+8(0.02) =1.08. $ A calculator gives approximately $1.082850933$. The small error shows the teaching of the note: near the base point, one linear model combines both input changes into a useful estimate. The graph of this affine approximation is the [[15B2 - Tangent Planes and Normal Vectors|tangent plane]]. > [!ML/AI] > Near parameters $\theta$, a loss obeys $L(\theta+\Delta)\approx L(\theta)+\nabla L(\theta)^\top\Delta$. Choosing $\Delta=-\eta\nabla L(\theta)$ predicts the decrease $-\eta\|\nabla L(\theta)\|^2$, which is the local justification for a gradient-descent step.