Gradient Descent Algorithm

Gradient Descent Algorithm - Main Topic

argmin over w: ‖X_w − y‖₂²
[ − x₁ − ]                       [ x₁ᵀw₁ ]
[ − x₂ − ] × [ w₁ w₂ … wₙ ] = [ x₂ᵀw₂ ]
[ −  ⋮ − ]                       [   ⋮   ]
[ − xₙ − ]                       [ xₙᵀwₙ ]

Gradient Descent

Let's consider the following optimization problem:

minimize over x: f(x) := x²
f(x) = f(xₜ) + ∇f(xₜ)ᵀ(x − xₜ) + α/2 + (x − xₜ)ᵀ(x − xₜ)
(x − xₜ)ᵀ(x − xₜ) = ‖x − xₜ‖₂²

or the L2 norm squared. Therefore,

f(x) = f(xₜ) + ∇f(xₜ)ᵀ(x − xₜ) + α/2 + ‖x − xₜ‖₂²

Now take the derivative of this formula.

0 + ∇f(xₜ) + (α/2) 2(x − xₜ) × I = 0
α(x − xₜ) = −∇f(xₜ)
x − xₜ = −(1/α)∇f(xₜ)
x = xₜ − (1/α)∇f(xₜ)