A Vision Transformer finishes with a sequence, not a single image vector. For ViT-B/16, 196 spatial patch rows leave the encoder. A classifier needs one stable row to read without flattening every patch or treating every location equally.
The CLS (classification) token supplies that row. It is a learned, non-spatial vector prepended at index 0, so the encoder receives CLS + patch₁ + … + patchₙ. Self-attention updates every row, including CLS. After the final block, a linear head reads the final CLS representation as the image-level summary.
Follow one CLS row from patches to logits
Use the example and layer controls, click any patch to inspect its score or attention weight, or play the complete seven-stage sequence. The same image-derived field persists from patchification through the CLS query, so the animation changes the mechanism rather than replacing it with a new diagram at every step.
What the encoder actually does
1. Prepend one learned row
Patch embedding produces N spatial vectors. CLS is an additional learned parameter with the same width D; it is not a patch and does not correspond to an image coordinate. Prepending it changes the sequence from N × D to (N+1) × D.
2. Add a position to every row
A learned position vector is added to each sequence row. CLS receives position 0, while patch rows receive positions 1…N. Positions are added, not concatenated, so every row keeps width D.
3. Let CLS participate in bidirectional self-attention
Inside a vanilla ViT encoder, CLS can query patch keys and patch rows can query CLS in the same bidirectional attention matrix. The displayed CLS row follows the familiar rule:
softmax(qCLS Kᵀ / √dₕ) V
The weighted value mix contributes to the CLS update, while residual paths preserve the incoming representation. Repeating this process through the encoder lets the CLS row accumulate image-level evidence.
4. Read only the final CLS row
For image classification, the final normalized CLS vector enters a linear classifier. The patch rows still exist and remain useful for dense or token-level tasks; the classifier simply chooses the designated CLS row as its readout.
What CLS does—and does not—promise
A CLS token is a learned alternative to pooling, not a magical compression primitive. It adds one token to every encoder block, so its cost includes the extra projection, attention, and MLP work throughout the stack. Its usefulness also depends on the training objective: the model must learn to place image-level information in that row.
The attention distribution for a CLS query is inspectable and can show where that query placed weight. It does not, by itself, prove which patches causally determined the prediction; explanations require stronger interventions or attribution checks.
| Readout strategy | Mechanism | Main trade-off |
|---|---|---|
| CLS token | Learn one extra row through every encoder block | Flexible learned readout, but adds a token and needs task training |
| Global average pooling | Average final patch rows | No special token, but every retained patch receives equal weight |
| Attention pooling | Use one or more learned queries after encoding | Flexible pooling with a separate query-to-source attention bill |
| Dense patch outputs | Keep spatial rows for a task head | Preserves local detail; no single image vector unless one is added |
Architectures reuse the idea differently. ViT reads CLS for supervised classification; DINO trains an image-level CLS objective; DeiT adds a separate distillation token beside CLS; dense systems such as ViTDet and segmentation heads retain patch outputs instead of treating CLS as the only useful representation.
Related concepts
Trace how local windows, shifted cross-window exchange, and patch merging turn one high-resolution token grid into a multi-scale vision hierarchy.
How multi-head attention runs scaled dot-product attention in parallel across several representation subspaces to build context-aware token embeddings.
Explore how positional embeddings enable Vision Transformers (ViT) to process sequential data by encoding relative positions.
Follow one image patch through Q/K/V projection, scaled scores, row-wise softmax, value mixing, and the residual update inside a Vision Transformer.
Learn adaptive tiling in vision transformers: dynamically partition images based on visual complexity to reduce token counts while preserving detail.
Learn ALiBi, the position encoding method that adds linear biases to attention scores for exceptional length extrapolation in transformers.
