The Attention Mechanism
How dynamic weighting allows models to focus on relevant context across long sequences.
The Attention Mechanism (Bahdanau et al., 2014; Vaswani et al., 2017) allows neural networks to dynamically weight and focus on relevant parts of an input sequence. Instead of compressing an entire sequence into a single static context vector, Attention computes dynamic alignment scores between Query, Key, and Value vectors. Scaled Dot Product Attention calculates Attention(Q, K, V) = Softmax( Q K^T / sqrt(d_k) ) V, forming the core foundation of modern Transformer language models.