Article
Gradient Descent Algorithm
Gradient Descent Algorithm - Main Topic
argmin over w: ‖X_w − y‖₂²[ − x₁ − ] [ x₁ᵀw₁ ]
[ − x₂ − ] × [ w₁ w₂ … wₙ ] = [ x₂ᵀw₂ ]
[ − ⋮ − ] [ ⋮ ]
[ − xₙ − ] [ xₙᵀwₙ ]Gradient Descent
Let's consider the following optimization problem:
minimize over x: f(x) := x²f(x) = f(xₜ) + ∇f(xₜ)ᵀ(x − xₜ) + α/2 + (x − xₜ)ᵀ(x − xₜ)(x − xₜ)ᵀ(x − xₜ) = ‖x − xₜ‖₂²or the L2 norm squared. Therefore,
f(x) = f(xₜ) + ∇f(xₜ)ᵀ(x − xₜ) + α/2 + ‖x − xₜ‖₂²Now take the derivative of this formula.
0 + ∇f(xₜ) + (α/2) 2(x − xₜ) × I = 0α(x − xₜ) = −∇f(xₜ)x − xₜ = −(1/α)∇f(xₜ)x = xₜ − (1/α)∇f(xₜ)