Signal & Noise
Menu

Transformer

The neural architecture behind modern LLMs: stacked self-attention and feed-forward layers processing all tokens in parallel. Introduced in 2017's 'Attention Is All You Need'.

Related terms

← Back to the full glossary