When you take the function
and start gradient descent at \(x_0 = (6, 6)\) with learning rate \(\eta = \frac{1}{2}\) it diverges.
Gradient descent
Gradient descent is an optimization rule which starts at a point \(x_0\) and then applies the update rule
where \(\eta\) is the step length (learning rate) and \(d_k\) is the direction.
The direction is
Example
In general:
In general, \(x_n = (1 - 8\eta) \cdot x_{n-1}\). You can clearly see that any learning rate \(\eta > \frac{1}{4}\) will diverge (with \(\eta = \frac{1}{4}\) it jumps back and forth between \((6, 6)\) and \((-6, -6)\)). For this example, the learning rate \(\eta = \frac{1}{8}\) would find the solution in one step and any \(\eta < \frac{1}{4}\) will converge to the global optimum.