Skip to main content

CLS Token in Vision Transformers

Trace how a learned CLS row joins image patches, gathers evidence through self-attention, and becomes the image-level classification readout.

A Vision Transformer finishes with a sequence, not a single image vector. For ViT-B/16, 196 spatial patch rows leave the encoder. A classifier needs one stable row to read without flattening every patch or treating every location equally.

The CLS (classification) token supplies that row. It is a learned, non-spatial vector prepended at index 0, so the encoder receives CLS + patch₁ + … + patchₙ. Self-attention updates every row, including CLS. After the final block, a linear head reads the final CLS representation as the image-level summary.

Follow one CLS row from patches to logits

Use the example and layer controls, click any patch to inspect its score or attention weight, or play the complete seven-stage sequence. The same image-derived field persists from patchification through the CLS query, so the animation changes the mechanism rather than replacing it with a new diagram at every step.

What the encoder actually does

1. Prepend one learned row

Patch embedding produces N spatial vectors. CLS is an additional learned parameter with the same width D; it is not a patch and does not correspond to an image coordinate. Prepending it changes the sequence from N × D to (N+1) × D.

2. Add a position to every row

A learned position vector is added to each sequence row. CLS receives position 0, while patch rows receive positions 1…N. Positions are added, not concatenated, so every row keeps width D.

3. Let CLS participate in bidirectional self-attention

Inside a vanilla ViT encoder, CLS can query patch keys and patch rows can query CLS in the same bidirectional attention matrix. The displayed CLS row follows the familiar rule:

softmax(qCLS Kᵀ / √dₕ) V

The weighted value mix contributes to the CLS update, while residual paths preserve the incoming representation. Repeating this process through the encoder lets the CLS row accumulate image-level evidence.

4. Read only the final CLS row

For image classification, the final normalized CLS vector enters a linear classifier. The patch rows still exist and remain useful for dense or token-level tasks; the classifier simply chooses the designated CLS row as its readout.

What CLS does—and does not—promise

A CLS token is a learned alternative to pooling, not a magical compression primitive. It adds one token to every encoder block, so its cost includes the extra projection, attention, and MLP work throughout the stack. Its usefulness also depends on the training objective: the model must learn to place image-level information in that row.

The attention distribution for a CLS query is inspectable and can show where that query placed weight. It does not, by itself, prove which patches causally determined the prediction; explanations require stronger interventions or attribution checks.

Readout strategyMechanismMain trade-off
CLS tokenLearn one extra row through every encoder blockFlexible learned readout, but adds a token and needs task training
Global average poolingAverage final patch rowsNo special token, but every retained patch receives equal weight
Attention poolingUse one or more learned queries after encodingFlexible pooling with a separate query-to-source attention bill
Dense patch outputsKeep spatial rows for a task headPreserves local detail; no single image vector unless one is added

Architectures reuse the idea differently. ViT reads CLS for supervised classification; DINO trains an image-level CLS objective; DeiT adds a separate distillation token beside CLS; dense systems such as ViTDet and segmentation heads retain patch outputs instead of treating CLS as the only useful representation.

If you found this explanation helpful, consider sharing it with others.

Mastodon