deep_learning
2 entries
Một chút về Transformers (Phần 2): Transformers
Phần hai trình bày kiến trúc Transformer qua self-attention, Query–Key–Value, multi-head attention, residual stream, layer normalization và cách song song hóa masked attention.
Một chút về Transformers (Phần 1): Attention, Attention, Attention
Phần đầu của chuỗi Transformers, giải thích từ distributional hypothesis và word embeddings đến RNN, encoder–decoder cho dịch máy, rồi dẫn vào cơ chế attention.