Before Q, K, and V: Reconstructing the Transformer
Transformer attention mechanisms emerged from RNN limitations, offering direct access to all past inputs. The Q, K, and V matrices naturally followed to break symmetry, reduce memory, and enable efficient GPU computation.