f(θ) = ½ θᵀHθθ⋆κ(H) ≫ 1gradient descentmomentum∇θ L(θ)
x ∈ ℝⁿhθMθ = { hθ(x) : x ∈ X }∂hθ / ∂z₁∂hθ / ∂z₂hθ(x₀)
HomeWorkBlogNoteResearchProjects

vi

3 entries

  1. Tags
  2. vi

Một chút về Transformers (Phần 2): Transformers

Jun 20, 2025

Phần hai trình bày kiến trúc Transformer qua self-attention, Query–Key–Value, multi-head attention, residual stream, layer normalization và cách song song hóa masked attention.

StableNonslop

Một chút về Transformers (Phần 1): Attention, Attention, Attention

Jun 19, 2025

Phần đầu của chuỗi Transformers, giải thích từ distributional hypothesis và word embeddings đến RNN, encoder–decoder cho dịch máy, rồi dẫn vào cơ chế attention.

StableNonslop

Về một vài câu hỏi mà mình gặp khi interview C++

Apr 29, 2025

Một vài câu hỏi phỏng vấn C++ về queue, circular buffer, độ phức tạp, chi phí cấp phát bộ nhớ và các tính năng Modern C++.

StableNonslop