Multivariable

Gradient and Directional Derivative: Formulas and Examples

Gradient and Directional Derivative: Formulas and Examples — CalculusCalc cover image

The directional derivative tells you how fast a function of several variables changes as you move in a chosen direction, and the gradient is the vector that makes computing it easy:

$$D_{\mathbf u}f = \nabla f\cdot\mathbf u$$

This guide defines both, derives the formula, works through six examples from routine to exam-level, and explains why the gradient always points uphill.

The gradient

For a function \( f(x, y) \), the gradient is the vector of its partial derivatives:

$$\nabla f = \left\langle\frac{\partial f}{\partial x},\ \frac{\partial f}{\partial y}\right\rangle$$

(In three variables, add a third component \( \frac{\partial f}{\partial z} \).)

The symbol \( \nabla \) is read “del” or “nabla,” and \( \nabla f \) is read “grad f.” The gradient is a vector-valued function: at every point of the plane it gives you a vector, so you can picture it as a field of arrows drawn over the contour map.

The directional derivative

The directional derivative is the rate of change of \( f \) when you move in the direction of a unit vector \( \mathbf{u} \):

$$D_{\mathbf u}f = \nabla f\cdot\mathbf u$$

If you’re given a non-unit direction \( \mathbf v \), normalize it first: \( \mathbf u = \frac{\mathbf v}{|\mathbf v|} \).

The answer is a single number, a slope, measured in units of \( f \) per unit of distance traveled in the plane.

Why the formula works

Start at \( (a, b) \) and walk in the direction \( \mathbf u = \langle u_1, u_2\rangle \). After traveling a distance \( t \) you’re at \( (a + tu_1,\ b + tu_2) \), so the height along your path is the one-variable function

$$g(t) = f(a + tu_1,\ b + tu_2)$$

The directional derivative is, by definition, \( g'(0) \): the rate of change along the path at the start. Both coordinates depend on \( t \), so the multivariable chain rule gives

$$g'(0) = f_x(a,b)\,u_1 + f_y(a,b)\,u_2$$

That sum is exactly the dot product \( \nabla f\cdot\mathbf u \). In words: moving \( u_1 \) units in \( x \) changes \( f \) at rate \( f_x \), moving \( u_2 \) units in \( y \) changes it at rate \( f_y \), and for small steps the two effects simply add.

This is also why \( \mathbf u \) must have length 1. The parameter \( t \) measures distance only when each step of \( t \) moves you one unit.

Example 1: the standard problem

\( f(x, y) = x^2 + 3xy + y^2 \) at \( (1, 1) \) in the direction of \( \mathbf v = \langle3, 4\rangle \).

Step 1: gradient. \( f_x = 2x + 3y \), \( f_y = 3x + 2y \). At \( (1,1) \):

$$\nabla f(1, 1) = \langle5,\ 5\rangle$$

Step 2: unit vector. \( |\mathbf v| = 5 \), so \( \mathbf u = \left\langle\frac35, \frac45\right\rangle \).

Step 3: dot product.

$$D_{\mathbf u}f = 5\cdot\frac35 + 5\cdot\frac45 = 3 + 4 = 7$$

Moving from \( (1,1) \) in that direction, \( f \) increases at 7 units per unit of distance.

Example 2: direction given by an angle

Find the directional derivative of \( f(x, y) = x^2y \) at \( (2, 1) \) in the direction that makes an angle of \( \frac\pi3 \) with the positive \( x \)-axis.

A direction given by an angle \( \theta \) is already a unit vector: \( \mathbf u = \langle\cos\theta, \sin\theta\rangle = \left\langle\frac12, \frac{\sqrt3}{2}\right\rangle \).

The gradient is \( \nabla f = \langle2xy,\ x^2\rangle \), which is \( \langle4, 4\rangle \) at \( (2, 1) \). Then

$$D_{\mathbf u}f = 4\cdot\frac12 + 4\cdot\frac{\sqrt3}{2} = 2 + 2\sqrt3 \approx 5.464$$

Example 3: three variables

Find the rate of change of \( f(x, y, z) = xyz \) at \( (1, 2, 3) \) in the direction of \( \mathbf v = \langle1, 2, 2\rangle \).

The gradient has three components: \( \nabla f = \langle yz,\ xz,\ xy\rangle \), so \( \nabla f(1, 2, 3) = \langle6, 3, 2\rangle \). The length of \( \mathbf v \) is \( \sqrt{1 + 4 + 4} = 3 \), so

$$D_{\mathbf u}f = \frac{6\cdot1 + 3\cdot2 + 2\cdot2}{3} = \frac{16}{3}$$

Dividing the final dot product by \( |\mathbf v| \) is the same as normalizing first, and often quicker.

What the gradient means geometrically

Since \( D_{\mathbf u}f = |\nabla f|\cos\theta \), where \( \theta \) is the angle between \( \nabla f \) and \( \mathbf u \):

  • Steepest ascent is in the direction of \( \nabla f \) (\( \theta = 0 \)), at rate \( |\nabla f| \). In Example 1, \( |\nabla f| = \sqrt{50} = 5\sqrt2 \approx 7.07 \).
  • Steepest descent is in the direction \( -\nabla f \).
  • Zero change when \( \mathbf u \perp \nabla f \): the gradient is perpendicular to level curves (contour lines).

That’s why hikers’ paths of steepest climb cross contour lines at right angles, and why “gradient descent” is how machine learning models minimize error.

Because \( \cos\theta \) ranges from \( -1 \) to \( 1 \), every directional derivative at a point lies between \( -|\nabla f| \) and \( |\nabla f| \). No direction can beat the gradient direction.

Example 4: maximum rate of change

For \( f(x, y) = xe^y \) at \( (2, 0) \), find the direction of fastest increase, the maximum rate, and the directions in which \( f \) doesn’t change.

The gradient is \( \nabla f = \langle e^y,\ xe^y\rangle \), so \( \nabla f(2, 0) = \langle1, 2\rangle \).

  • Direction of fastest increase: \( \langle1, 2\rangle \), or as a unit vector \( \frac{1}{\sqrt5}\langle1, 2\rangle \).
  • Maximum rate: \( |\nabla f| = \sqrt5 \approx 2.236 \).
  • Fastest decrease: direction \( \langle-1, -2\rangle \), at rate \( -\sqrt5 \).
  • No change: perpendicular to \( \langle1, 2\rangle \), so \( \pm\frac{1}{\sqrt5}\langle2, -1\rangle \). Check: \( 1\cdot2 + 2\cdot(-1) = 0 \).

Example 5: a heat-seeking path

The temperature on a plate is \( T(x, y) = 100 - x^2 - 2y^2 \). A small bug at \( (3, 1) \) wants to warm up as fast as possible. Which way should it go?

\( \nabla T = \langle-2x,\ -4y\rangle \), so \( \nabla T(3, 1) = \langle-6, -4\rangle \). The bug should head in the direction \( \langle-6, -4\rangle \), toward the hot center, and the temperature initially rises at \( \sqrt{36 + 16} = 2\sqrt{13} \approx 7.21 \) degrees per unit of distance.

Notice that this direction doesn’t point straight at the origin, the hottest spot. The gradient gives the best direction right now, based on local slopes, not a route to the global maximum.

Example 6: gradients are normal to level curves

The circle \( x^2 + y^2 = 25 \) is a level curve of \( f = x^2 + y^2 \). At \( (3, 4) \), \( \nabla f = \langle6, 8\rangle \), which points straight out from the center, perpendicular to the circle.

Because the gradient is a normal vector, the tangent line at \( (3, 4) \) is \( 6(x - 3) + 8(y - 4) = 0 \), or \( 3x + 4y = 25 \). This is often the fastest way to find a tangent line to an implicit curve, and the same idea in three variables gives tangent planes to level surfaces.

Special cases

  • \( \mathbf u = \langle1, 0\rangle \) gives \( D_{\mathbf u}f = f_x \).
  • \( \mathbf u = \langle0, 1\rangle \) gives \( D_{\mathbf u}f = f_y \).

So partial derivatives are just directional derivatives along the axes.

Using the gradient for tangent planes

The tangent plane to \( z = f(x, y) \) at \( (a, b) \) is

$$z = f(a,b) + f_x(a,b)(x - a) + f_y(a,b)(y - b)$$

the two-variable version of the tangent line. It is also the linear approximation of \( f \) near \( (a, b) \): the change in \( f \) for a small step \( \Delta\mathbf r \) is about \( \nabla f\cdot\Delta\mathbf r \).

Common mistakes

  • Not normalizing the direction. Using \( \langle3, 4\rangle \) instead of \( \langle\frac35, \frac45\rangle \) in Example 1 gives 35 instead of 7. Always divide by the length.
  • Using a point instead of a direction. “Toward the point \( (4, 5) \) from \( (1, 1) \)” means \( \mathbf v = \langle4 - 1,\ 5 - 1\rangle = \langle3, 4\rangle \), not \( \langle4, 5\rangle \).
  • Forgetting to evaluate the gradient. The gradient formula has variables in it. Plug in the point before taking the dot product, or your answer won’t be a number.
  • Confusing the gradient with the directional derivative. The gradient is a vector; the directional derivative is a scalar. “Maximum rate of change” asks for \( |\nabla f| \), a number, while “direction of maximum increase” asks for the vector.
  • Assuming the gradient points to the maximum. It points in the locally steepest direction, which is generally not the direction of the highest point of the surface, as Example 5 showed.

Where it’s used

  • Optimization: the gradient is zero at critical points, the first step in finding critical points of surfaces. Lagrange multipliers extend the optimization toolkit by setting \( \nabla f \) parallel to \( \nabla g \).
  • Machine learning: gradient descent repeatedly steps in the direction \( -\nabla f \) to reduce an error function with thousands of variables.
  • Physics: heat flows in the direction \( -\nabla T \), and a force is minus the gradient of potential energy.
  • Maps: on a topographic map, gradient arrows cross contour lines at right angles, and their length shows how steep the terrain is.

Gradient and directional derivative calculator

Multivariable#22

Gradient & Directional Derivative

\(\nabla f\) and \(D_{\mathbf{u}} f = \nabla f \cdot \hat{\mathbf{u}}\) at a point.

For more examples, including three-variable functions, open the full gradient calculator. To check the individual components, the partial derivative calculator lists every first and second partial.

Practice problems

Try these before checking the answers.

  1. \( \nabla(x^2y) \text{ at } (1, 3) \)
  2. \( D_{\mathbf u}(x^2 + y^2) \text{ at } (1,1),\ \mathbf u = \langle\tfrac35, \tfrac45\rangle \)
  3. Max rate of change of \( xy \) at \( (3, 4) \)
  4. \( D_{\mathbf u}(\sin x + \cos y) \) at \( (0, 0) \) in the direction \( \langle1, 1\rangle \)
  5. The direction of steepest descent of \( x^2 + y^2 \) at \( (1, 2) \)
  6. The unit vectors in which \( xy \) has zero rate of change at \( (3, 4) \)

Answers: (1) \( \langle 6,\ 1\rangle \); (2) \( \frac{14}{5} \); (3) \( 5 \); (4) \( \frac{\sqrt2}{2} \); (5) \( \langle-2, -4\rangle \), or the unit vector \( \frac{1}{\sqrt5}\langle-1, -2\rangle \); (6) \( \pm\left\langle\frac35, -\frac45\right\rangle \).

For problem 4, \( \nabla f = \langle\cos x, -\sin y\rangle = \langle1, 0\rangle \) at the origin, and \( \mathbf u = \frac{1}{\sqrt2}\langle1, 1\rangle \). For problem 6, \( \nabla(xy) = \langle4, 3\rangle \) at \( (3, 4) \), and \( \langle3, -4\rangle \) is perpendicular to it.

FAQ

Why must the direction be a unit vector?

Otherwise the answer is scaled by the vector’s length and no longer measures change per unit distance.

Where is the gradient zero?

At critical points of \( f \): candidates for maxima, minima and saddle points.

Can a directional derivative be negative?

Yes. A negative value means \( f \) decreases as you move in that direction. It happens whenever the direction makes an angle of more than 90° with the gradient.

Is the gradient a vector or a number?

A vector. Its direction is the direction of steepest increase and its length is the steepest rate. The directional derivative, by contrast, is a number.

What is the difference between the gradient and the derivative?

For a function of one variable, the derivative is a single slope. With several variables there is a slope in every direction, and the gradient packages all of them: any directional slope is the gradient dotted with the direction.

Why is the gradient perpendicular to level curves?

Along a level curve, \( f \) is constant, so its rate of change in the tangent direction is zero. That means \( \nabla f\cdot\mathbf u = 0 \) for the tangent vector, which is exactly the statement that they’re perpendicular.

Further reading

Calculators for this topic

Leave a comment

Your email address will not be published. Required fields are marked *