f(θ) = ½ θᵀHθθ⋆κ(H) ≫ 1gradient descentmomentum∇θ L(θ)
x ∈ ℝⁿhθMθ = { hθ(x) : x ∈ X }∂hθ / ∂z₁∂hθ / ∂z₂hθ(x₀)
HomeWorkBlogNoteResearchProjects

deep_learning

2 entries

  1. Tags
  2. deep_learning

Một chút về Transformers (Phần 2): Transformers

Jun 20, 2025

Phần hai trình bày kiến trúc Transformer qua self-attention, Query–Key–Value, multi-head attention, residual stream, layer normalization và cách song song hóa masked attention.

StableNonslop

Một chút về Transformers (Phần 1): Attention, Attention, Attention

Jun 19, 2025

Phần đầu của chuỗi Transformers, giải thích từ distributional hypothesis và word embeddings đến RNN, encoder–decoder cho dịch máy, rồi dẫn vào cơ chế attention.

StableNonslop