f(θ) = ½ θᵀHθ
θ⋆
κ(H) ≫ 1
gradient descent
momentum
∇θ L(θ)
x ∈ ℝⁿ
hθ
Mθ = { hθ(x) : x ∈ X }
∂hθ / ∂z₁
∂hθ / ∂z₂
hθ(x₀)
Home
Work
Blog
Note
Research
Projects
deep_learning
2 entries
Tags
deep_learning
Một chút về Transformers (Phần 2): Transformers
Jun 20, 2025
Stable
Nonslop
Một chút về Transformers (Phần 1): Attention, Attention, Attention
Jun 19, 2025
Stable
Nonslop