The directional derivative tells you how fast a function of several variables changes as you move in a chosen direction, and the gradient is the vector that makes computing it easy:
$$D_{\mathbf u}f = \nabla f\cdot\mathbf u$$
This guide defines both, derives the formula, works through six examples from routine to exam-level, and explains why the gradient always points uphill.
The gradient
For a function \( f(x, y) \), the gradient is the vector of its partial derivatives:
$$\nabla f = \left\langle\frac{\partial f}{\partial x},\ \frac{\partial f}{\partial y}\right\rangle$$
(In three variables, add a third component \( \frac{\partial f}{\partial z} \).)
The symbol \( \nabla \) is read “del” or “nabla,” and \( \nabla f \) is read “grad f.” The gradient is a vector-valued function: at every point of the plane it gives you a vector, so you can picture it as a field of arrows drawn over the contour map.
The directional derivative
The directional derivative is the rate of change of \( f \) when you move in the direction of a unit vector \( \mathbf{u} \):
$$D_{\mathbf u}f = \nabla f\cdot\mathbf u$$
If you’re given a non-unit direction \( \mathbf v \), normalize it first: \( \mathbf u = \frac{\mathbf v}{|\mathbf v|} \).
The answer is a single number, a slope, measured in units of \( f \) per unit of distance traveled in the plane.
Why the formula works
Start at \( (a, b) \) and walk in the direction \( \mathbf u = \langle u_1, u_2\rangle \). After traveling a distance \( t \) you’re at \( (a + tu_1,\ b + tu_2) \), so the height along your path is the one-variable function
$$g(t) = f(a + tu_1,\ b + tu_2)$$
The directional derivative is, by definition, \( g'(0) \): the rate of change along the path at the start. Both coordinates depend on \( t \), so the multivariable chain rule gives
$$g'(0) = f_x(a,b)\,u_1 + f_y(a,b)\,u_2$$
That sum is exactly the dot product \( \nabla f\cdot\mathbf u \). In words: moving \( u_1 \) units in \( x \) changes \( f \) at rate \( f_x \), moving \( u_2 \) units in \( y \) changes it at rate \( f_y \), and for small steps the two effects simply add.
This is also why \( \mathbf u \) must have length 1. The parameter \( t \) measures distance only when each step of \( t \) moves you one unit.
Example 1: the standard problem
\( f(x, y) = x^2 + 3xy + y^2 \) at \( (1, 1) \) in the direction of \( \mathbf v = \langle3, 4\rangle \).
Step 1: gradient. \( f_x = 2x + 3y \), \( f_y = 3x + 2y \). At \( (1,1) \):
$$\nabla f(1, 1) = \langle5,\ 5\rangle$$
Step 2: unit vector. \( |\mathbf v| = 5 \), so \( \mathbf u = \left\langle\frac35, \frac45\right\rangle \).
Step 3: dot product.
$$D_{\mathbf u}f = 5\cdot\frac35 + 5\cdot\frac45 = 3 + 4 = 7$$
Moving from \( (1,1) \) in that direction, \( f \) increases at 7 units per unit of distance.
Example 2: direction given by an angle
Find the directional derivative of \( f(x, y) = x^2y \) at \( (2, 1) \) in the direction that makes an angle of \( \frac\pi3 \) with the positive \( x \)-axis.
A direction given by an angle \( \theta \) is already a unit vector: \( \mathbf u = \langle\cos\theta, \sin\theta\rangle = \left\langle\frac12, \frac{\sqrt3}{2}\right\rangle \).
The gradient is \( \nabla f = \langle2xy,\ x^2\rangle \), which is \( \langle4, 4\rangle \) at \( (2, 1) \). Then
$$D_{\mathbf u}f = 4\cdot\frac12 + 4\cdot\frac{\sqrt3}{2} = 2 + 2\sqrt3 \approx 5.464$$
Example 3: three variables
Find the rate of change of \( f(x, y, z) = xyz \) at \( (1, 2, 3) \) in the direction of \( \mathbf v = \langle1, 2, 2\rangle \).
The gradient has three components: \( \nabla f = \langle yz,\ xz,\ xy\rangle \), so \( \nabla f(1, 2, 3) = \langle6, 3, 2\rangle \). The length of \( \mathbf v \) is \( \sqrt{1 + 4 + 4} = 3 \), so
$$D_{\mathbf u}f = \frac{6\cdot1 + 3\cdot2 + 2\cdot2}{3} = \frac{16}{3}$$
Dividing the final dot product by \( |\mathbf v| \) is the same as normalizing first, and often quicker.
What the gradient means geometrically
Since \( D_{\mathbf u}f = |\nabla f|\cos\theta \), where \( \theta \) is the angle between \( \nabla f \) and \( \mathbf u \):
- Steepest ascent is in the direction of \( \nabla f \) (\( \theta = 0 \)), at rate \( |\nabla f| \). In Example 1, \( |\nabla f| = \sqrt{50} = 5\sqrt2 \approx 7.07 \).
- Steepest descent is in the direction \( -\nabla f \).
- Zero change when \( \mathbf u \perp \nabla f \): the gradient is perpendicular to level curves (contour lines).
That’s why hikers’ paths of steepest climb cross contour lines at right angles, and why “gradient descent” is how machine learning models minimize error.
Because \( \cos\theta \) ranges from \( -1 \) to \( 1 \), every directional derivative at a point lies between \( -|\nabla f| \) and \( |\nabla f| \). No direction can beat the gradient direction.
Example 4: maximum rate of change
For \( f(x, y) = xe^y \) at \( (2, 0) \), find the direction of fastest increase, the maximum rate, and the directions in which \( f \) doesn’t change.
The gradient is \( \nabla f = \langle e^y,\ xe^y\rangle \), so \( \nabla f(2, 0) = \langle1, 2\rangle \).
- Direction of fastest increase: \( \langle1, 2\rangle \), or as a unit vector \( \frac{1}{\sqrt5}\langle1, 2\rangle \).
- Maximum rate: \( |\nabla f| = \sqrt5 \approx 2.236 \).
- Fastest decrease: direction \( \langle-1, -2\rangle \), at rate \( -\sqrt5 \).
- No change: perpendicular to \( \langle1, 2\rangle \), so \( \pm\frac{1}{\sqrt5}\langle2, -1\rangle \). Check: \( 1\cdot2 + 2\cdot(-1) = 0 \).
Example 5: a heat-seeking path
The temperature on a plate is \( T(x, y) = 100 - x^2 - 2y^2 \). A small bug at \( (3, 1) \) wants to warm up as fast as possible. Which way should it go?
\( \nabla T = \langle-2x,\ -4y\rangle \), so \( \nabla T(3, 1) = \langle-6, -4\rangle \). The bug should head in the direction \( \langle-6, -4\rangle \), toward the hot center, and the temperature initially rises at \( \sqrt{36 + 16} = 2\sqrt{13} \approx 7.21 \) degrees per unit of distance.
Notice that this direction doesn’t point straight at the origin, the hottest spot. The gradient gives the best direction right now, based on local slopes, not a route to the global maximum.
Example 6: gradients are normal to level curves
The circle \( x^2 + y^2 = 25 \) is a level curve of \( f = x^2 + y^2 \). At \( (3, 4) \), \( \nabla f = \langle6, 8\rangle \), which points straight out from the center, perpendicular to the circle.
Because the gradient is a normal vector, the tangent line at \( (3, 4) \) is \( 6(x - 3) + 8(y - 4) = 0 \), or \( 3x + 4y = 25 \). This is often the fastest way to find a tangent line to an implicit curve, and the same idea in three variables gives tangent planes to level surfaces.
Special cases
- \( \mathbf u = \langle1, 0\rangle \) gives \( D_{\mathbf u}f = f_x \).
- \( \mathbf u = \langle0, 1\rangle \) gives \( D_{\mathbf u}f = f_y \).
So partial derivatives are just directional derivatives along the axes.
Using the gradient for tangent planes
The tangent plane to \( z = f(x, y) \) at \( (a, b) \) is
$$z = f(a,b) + f_x(a,b)(x - a) + f_y(a,b)(y - b)$$
the two-variable version of the tangent line. It is also the linear approximation of \( f \) near \( (a, b) \): the change in \( f \) for a small step \( \Delta\mathbf r \) is about \( \nabla f\cdot\Delta\mathbf r \).
Common mistakes
- Not normalizing the direction. Using \( \langle3, 4\rangle \) instead of \( \langle\frac35, \frac45\rangle \) in Example 1 gives 35 instead of 7. Always divide by the length.
- Using a point instead of a direction. “Toward the point \( (4, 5) \) from \( (1, 1) \)” means \( \mathbf v = \langle4 - 1,\ 5 - 1\rangle = \langle3, 4\rangle \), not \( \langle4, 5\rangle \).
- Forgetting to evaluate the gradient. The gradient formula has variables in it. Plug in the point before taking the dot product, or your answer won’t be a number.
- Confusing the gradient with the directional derivative. The gradient is a vector; the directional derivative is a scalar. “Maximum rate of change” asks for \( |\nabla f| \), a number, while “direction of maximum increase” asks for the vector.
- Assuming the gradient points to the maximum. It points in the locally steepest direction, which is generally not the direction of the highest point of the surface, as Example 5 showed.
Where it’s used
- Optimization: the gradient is zero at critical points, the first step in finding critical points of surfaces. Lagrange multipliers extend the optimization toolkit by setting \( \nabla f \) parallel to \( \nabla g \).
- Machine learning: gradient descent repeatedly steps in the direction \( -\nabla f \) to reduce an error function with thousands of variables.
- Physics: heat flows in the direction \( -\nabla T \), and a force is minus the gradient of potential energy.
- Maps: on a topographic map, gradient arrows cross contour lines at right angles, and their length shows how steep the terrain is.
Gradient and directional derivative calculator
Gradient & Directional Derivative
\(\nabla f\) and \(D_{\mathbf{u}} f = \nabla f \cdot \hat{\mathbf{u}}\) at a point.
For more examples, including three-variable functions, open the full gradient calculator. To check the individual components, the partial derivative calculator lists every first and second partial.
Practice problems
Try these before checking the answers.
- \( \nabla(x^2y) \text{ at } (1, 3) \)
- \( D_{\mathbf u}(x^2 + y^2) \text{ at } (1,1),\ \mathbf u = \langle\tfrac35, \tfrac45\rangle \)
- Max rate of change of \( xy \) at \( (3, 4) \)
- \( D_{\mathbf u}(\sin x + \cos y) \) at \( (0, 0) \) in the direction \( \langle1, 1\rangle \)
- The direction of steepest descent of \( x^2 + y^2 \) at \( (1, 2) \)
- The unit vectors in which \( xy \) has zero rate of change at \( (3, 4) \)
Answers: (1) \( \langle 6,\ 1\rangle \); (2) \( \frac{14}{5} \); (3) \( 5 \); (4) \( \frac{\sqrt2}{2} \); (5) \( \langle-2, -4\rangle \), or the unit vector \( \frac{1}{\sqrt5}\langle-1, -2\rangle \); (6) \( \pm\left\langle\frac35, -\frac45\right\rangle \).
For problem 4, \( \nabla f = \langle\cos x, -\sin y\rangle = \langle1, 0\rangle \) at the origin, and \( \mathbf u = \frac{1}{\sqrt2}\langle1, 1\rangle \). For problem 6, \( \nabla(xy) = \langle4, 3\rangle \) at \( (3, 4) \), and \( \langle3, -4\rangle \) is perpendicular to it.
FAQ
Why must the direction be a unit vector?
Otherwise the answer is scaled by the vector’s length and no longer measures change per unit distance.
Where is the gradient zero?
At critical points of \( f \): candidates for maxima, minima and saddle points.
Can a directional derivative be negative?
Yes. A negative value means \( f \) decreases as you move in that direction. It happens whenever the direction makes an angle of more than 90° with the gradient.
Is the gradient a vector or a number?
A vector. Its direction is the direction of steepest increase and its length is the steepest rate. The directional derivative, by contrast, is a number.
What is the difference between the gradient and the derivative?
For a function of one variable, the derivative is a single slope. With several variables there is a slope in every direction, and the gradient packages all of them: any directional slope is the gradient dotted with the direction.
Why is the gradient perpendicular to level curves?
Along a level curve, \( f \) is constant, so its rate of change in the tangent direction is zero. That means \( \nabla f\cdot\mathbf u = 0 \) for the tangent vector, which is exactly the statement that they’re perpendicular.
Further reading
- Directional Derivatives (Paul’s Online Math Notes) — the formal definition, a proof of the dot-product formula and more examples in two and three variables.
- Directional Derivatives and the Gradient (OpenStax Calculus Volume 3) — a textbook treatment with contour-map illustrations and exercises.
Calculators for this topic
Keep learning
Double Integrals: How to Evaluate Them Step by Step
Evaluate double integrals as iterated integrals, switch the order of integration, and use polar coordinates. Step-by-step worked examples.
Partial Derivatives: How to Calculate Them (with Examples)
A partial derivative differentiates in one variable, holding the others constant. Notation, examples, second partials and Clairaut’s theorem.

