f(θ) = ½ θᵀHθθ⋆κ(H) ≫ 1gradient descentmomentum∇θ L(θ)
x ∈ ℝⁿhθMθ = { hθ(x) : x ∈ X }∂hθ / ∂z₁∂hθ / ∂z₂hθ(x₀)
HomeWorkBlogNoteResearchProjects

reinforcement_learning

1 entries

  1. Tags
  2. reinforcement_learning

Note on Sutton & Barto — Chapter 3: Finite Markov Decision Process (MDP)

Jul 22, 2026
StableNearly slop