Transformers, from a distance — study guide
The concept's fragments, read in order.
The architecture behind modern models
A transformer is a neural network architecture — a particular design for a model that learns its behavior from data. It is the design behind today's language models, and behind a great many models that work on other kinds of input as well. Naming the architecture is the first step to seeing why those models act the way they do.
What a transformer does is take a sequence of tokens — the small pieces an input is broken into — and turn it into context-aware representations, one per position. A representation is a vector, a list of numbers in the sense that embeddings already use, and "context-aware" is the whole trick: each piece is described not on its own but in light of the pieces around it. The same word carries a different representation depending on the company it keeps.
This concept stays at the level of shapes and moves, not equations. The point is to see how a sequence goes in, how the pieces come to inform each other, and how the result comes out richer than it went in — enough to understand what the models built this way are actually doing.
From text to vectors
Before a transformer can do anything, the input has to become numbers. Text is first split into tokens — word pieces, not always whole words like breaking a sentence into small snap-together blocks the model already knows, rather than swallowing it whole. Common words may be a single token, while a rare or long word gets chopped into a few standard pieces the model has seen before. The split is fixed and mechanical, and it is how the model handles anything, including words it has never met.
Each token then becomes an embedding vector: a list of numbers positioned so that pieces with related meanings sit near each other. This is embeddings doing their usual job — turning an item into coordinates a machine can compare — applied here to the tokens of a sequence. What the model receives is not the sentence but a row of these vectors, one per token, lined up in order.
That row is the transformer's actual input. Everything the architecture does from here is arithmetic over those vectors, which is why the tokenizing-and-embedding step has to happen first.
Weighing what matters
Attention is the central mechanism of a transformer — the part that makes the representations context-aware. For each token, the model weighs how much every other token in the sequence matters to it, and blends those tokens together according to those weights. A word like "it" leans heavily on whatever noun it refers to and barely on the rest like reading one word while glancing back at the other words that give it meaning, leaning on the ones that matter most. The weights are not fixed; the model learned how to assign them from data.
Because a token attends to the other tokens directly, it can draw on anything in the sequence — a word right beside it or one far away — without the meaning having to be passed along hand to hand. The output for each position is a new vector: the same token, now carrying a description shaped by the parts of the sequence that actually bear on it. This is often called self-attention, because the sequence is attending to itself.
That single idea is what the original architecture leaned on so completely that the paper introducing it was titled "Attention Is All You Need." Strip away the surrounding machinery and this weighing of one token against the rest is what remains.
All at once, any distance
Earlier sequence models read one step at a time, carrying a running summary forward from token to token. That order was a straitjacket: a word near the end could only feel a word near the start through everything in between, and the reading could not be split up because each step waited on the one before it.
A transformer drops the step-by-step reading. It processes all positions in parallel, and attention lets any token relate to any other token directly, however far apart they sit. Nothing has to be relayed through the tokens in between, so a long-range connection is no harder to make than a neighboring one.
Doing the work all at once instead of in sequence is also what lets these models train efficiently on the enormous amounts of text they are built from. The architecture fits the hardware, and that fit is a large part of why it scales.
Refining, layer by layer
A transformer is not one pass but many. It stacks identical layers, and each layer has the same two parts: an attention step, where the tokens weigh and blend each other, followed by a feed-forward step that processes each position a little further on its own. The representations flow up through the stack, reworked at every layer like a draft passing through a series of editors, each pass refining the understanding a little further.
Stacking is what lets the model build up structure. An early layer might settle plain, local relationships, while a later layer, working over representations that already carry context, can capture something more abstract. Each pass has more to work with than the pass beneath it, so depth compounds rather than merely repeats.
The result that leaves the top of the stack is the context-aware representation the whole architecture was after — the same tokens that went in, now described with everything the layers managed to fold into them.
Why this architecture took over
The transformer was introduced in 2017, in a paper by Vaswani and colleagues titled "Attention Is All You Need." It did not arrive into an empty field; it displaced the sequential architectures that had held the ground before it, and it did so quickly.
Three advantages stacked up. Processing all positions in parallel meant the models could train far faster and on far more data than the step-by-step designs allowed. Attention gave every token a direct line to every other, so long-range context that older models struggled to hold came almost for free. And the design kept paying off as it grew — more data and more compute reliably bought a better model, with no obvious wall.
That combination is why nearly everything in modern language modeling is built on this one architecture. It was not a small improvement on what came before so much as a different footing, and the field moved onto it.
Not just language
Nothing in the architecture is really about words. A transformer takes a sequence of tokens and relates them by attention; what those tokens stand for is a detail of how the input was prepared, not of the machinery that processes it. Break the input into pieces, embed each piece as a vector, and the same design applies.
So the same architecture handles images, cut into patches instead of words, and audio, sliced into short segments. It handles mixed inputs too, where text and pictures share one sequence and can attend to each other across the boundary. The task type changes, but the engine does not.
This is why the transformer is better thought of as a general engine for sequences of tokens than as a language tool that happens to generalize. Language was where it landed first, not the limit of what it does.
What is under the models you call
The models an agent pipeline calls out to are transformers. Whatever the wrapper around them looks like, underneath is this architecture: input broken into tokens, each token embedded as a vector, and layers of attention turning that sequence into context-aware representations.
Holding that picture demystifies a fair amount. A model does not read your input as a paragraph; it reads a sequence of tokens, and it works by weighing those tokens against each other rather than by following rules someone wrote down. Its strengths and its limits both trace back to that — to what a stack of attention layers over a row of token vectors can and cannot do.
Knowing the architecture will not tell you everything about how such a model behaves, but it tells you what kind of thing you are dealing with. The rest of that story builds on this foundation, not around it.