Blog
Why do we need KV caching?
An explanation of causal attention's repeated work during decoding and how reusing keys and values makes autoregressive generation cheaper in compute at a memory cost.
Why I want to learn to code again?
A personal return to hands-on programming through a ray tracer built from scratch in Rust, with detours into Rust's safety model and computer graphics.
Một chút về Transformers (Phần 2): Transformers
Phần hai trình bày kiến trúc Transformer qua self-attention, Query–Key–Value, multi-head attention, residual stream, layer normalization và cách song song hóa masked attention.
Một chút về Transformers (Phần 1): Attention, Attention, Attention
Phần đầu của chuỗi Transformers, giải thích từ distributional hypothesis và word embeddings đến RNN, encoder–decoder cho dịch máy, rồi dẫn vào cơ chế attention.
Về một vài câu hỏi mà mình gặp khi interview C++
Một vài câu hỏi phỏng vấn C++ về queue, circular buffer, độ phức tạp, chi phí cấp phát bộ nhớ và các tính năng Modern C++.