Skip to main content

Vision-Language Alignment: CLIP, BLIP-2, LLaVA, Flamingo

Summary
CLIP, BLIP-2, LLaVA and Flamingo align images with text by different routes. Compare their objectives, bridges, data, and which parts train or stay frozen.

Every vision-language model has to bring images and text to a place where they can be compared or combined. CLIP does it by training two encoders into one shared embedding space. BLIP-2, LLaVA and Flamingo start instead from a pretrained vision encoder and a pretrained language model, and train a bridge that lets the language model read visual features. This page compares the four mechanisms, what each one trains and freezes, and which tasks each one suits. The alignment problem page covers the objective in general, and each model below links to its own paper walkthrough.

TL;DR

  • All four models align visual and textual representations. They differ in where the alignment lives and in which weights learn it.
  • CLIP (Radford et al., 2021) trains an image encoder and a text encoder from scratch with a symmetric contrastive loss on 400M image-text pairs. Matching images and captions land close together in one embedding space, which gives retrieval and zero-shot classification but no text generation.
  • BLIP-2 (Li et al., 2023) keeps a frozen image encoder and a frozen LLM and trains a lightweight Querying Transformer (Q-Former) between them, whose 32 learned queries pull visual features out of the encoder. Stage 1 aligns the Q-Former with text; stage 2 connects it to the LLM through a linear projection.
  • LLaVA (Liu et al., 2023) projects the patch features of a frozen CLIP ViT-L/14 into the LLM's word-embedding space, with a linear layer in the original and a two-layer MLP in LLaVA-1.5. Stage 1 trains only the projector; stage 2 also fine-tunes the LLM.
  • Flamingo (Alayrac et al., 2022) keeps a frozen vision encoder and a frozen language model and trains a Perceiver Resampler plus new tanh-gated cross-attention layers between the LM blocks. It reads interleaved images and text and learns new tasks from a few examples in its prompt.

The fundamental design choice

Train a joint embedding space

Two encoders, one per modality, each end in a projection to the same vector space. Training pulls matched image-caption pairs together and pushes mismatched pairs apart, so the alignment is geometric: the similarity of an image and a text is the dot product of two unit vectors.

Images and texts are encoded independently, so embeddings can be computed once and indexed, and a classifier is simply a set of text prompts. The cost is that each image and each caption becomes one vector, compared by one score. Nothing in this setup writes text, and a single vector drops detail that a question about the image might need. Contrastive training also leaves the two modalities in separate regions of the shared space, the modality gap.

Project a vision encoder into a language model

The second family keeps a pretrained LLM as the part that reasons and writes, and turns the image into something it can read. Alignment here means the language model can use the visual features. The bridge is trained with the language model's own next-token loss, so it is judged by whether the text that follows the image becomes more likely. This gives captioning, question answering and dialogue, but every query runs the language model, so the result is not a retrieval index.

Three further choices

Within the projection family, three choices separate BLIP-2, LLaVA and Flamingo.

  • Where visual information enters. BLIP-2 and LLaVA place visual tokens in the LLM's input sequence, where its existing self-attention reads them like text, with no new layers. Flamingo keeps the input sequence as text and adds new cross-attention layers inside the LM, through which text tokens read the visual tokens.
  • Whether the LLM learns. BLIP-2 and Flamingo keep the LLM frozen, which preserves its text abilities and keeps the trainable part small; BLIP-2 reports 54 times fewer trainable parameters than Flamingo-80B. LLaVA fine-tunes the LLM in its second stage, giving up that protection for stronger instruction following. Parameter-efficient methods such as LoRA sit between the two (vision-language adapters).
  • How many visual tokens. LLaVA passes every patch through: 256 tokens for a 224-pixel image with ViT-L/14, and 576 at 336 pixels in LLaVA-1.5. BLIP-2 compresses any image to 32 query outputs, and Flamingo's resampler to 64 tokens per image. Fewer tokens use less compute and context; more tokens keep more spatial detail.

The two families are stages of one pipeline more than rival lineages. All three bridge models start from a vision encoder trained with a contrastive image-text loss: BLIP-2 uses CLIP's ViT-L/14 or EVA-CLIP's ViT-g/14, LLaVA uses CLIP's ViT-L/14, and Flamingo pre-trains its own NFNet-F6 contrastively before freezing it.

Four architectures, stage by stage

Each panel follows one model from its input to its training signal. Pick a stage to see which parts learn in it; every part is marked trainable, frozen, or not used, in text and with an icon.

What trains in each stage: CLIP, BLIP-2, LLaVA and Flamingo
Training stage

Stage 1: BLIP-2 trains only its Q-Former, with no LLM in the loop, and LLaVA trains only its projector. CLIP and Flamingo train in one stage, so they read the same in both.

CLIP, one stage: trains the image encoder, text encoder and linear projections. BLIP-2, stage 1: trains the Q-Former; frozen: image encoder; not used: linear projection and LLM. LLaVA, stage 1: trains the projector; frozen: vision encoder and LLM. Flamingo, one stage: trains the Perceiver Resampler and gated cross-attention; frozen: vision encoder and LM blocks.

  • CLIP Radford et al., 2021

    Contrastive dual encoder

    One stage: contrastive pre-training

    1. Image

      from a web image-text pair

      Caption

      its paired text

    2. Image encoderTrainable

      ResNet or ViT, from scratch

      Text encoderTrainable

      Transformer, from scratch

    3. Linear projectionsTrainable

      both towers into one shared embedding space

    4. Symmetric contrastive loss

      over the N × N image-text similarities; matched pairs pulled together

    CLIP paper walkthrough
  • BLIP-2 Li et al., 2023

    Q-Former between two frozen models

    Stage 1: representation learning

    1. Image

      with a caption

      Caption

      read by the Q-Former

    2. Image encoderFrozen

      ViT-L/14 (CLIP) or ViT-g/14 (EVA-CLIP)

    3. Q-FormerTrainable

      32 learned queries cross-attend to image features

    4. Linear projectionNot used

      32 query outputs into the LLM input space

    5. LLMNot used

      OPT or FlanT5

    6. Training signal

      ITC, ITM and ITG losses on the Q-Former

    BLIP-2 paper walkthrough
  • LLaVA Liu et al., 2023

    Projection into the LLM input

    Stage 1: feature alignment

    1. Image

      one per conversation

      Text

      caption as a brief description

    2. Vision encoderFrozen

      CLIP ViT-L/14 patch features

    3. ProjectorTrainable

      linear (LLaVA) or two-layer MLP (LLaVA-1.5)

    4. LLMFrozen

      Vicuna reads visual and text tokens together

    5. Training signal

      next-token loss on 595K CC3M captions

    LLaVA paper walkthrough
  • Flamingo Alayrac et al., 2022

    Gated cross-attention into a frozen LM

    One stage: training on a multimodal web mixture

    1. Interleaved images and text

      web pages, image-text and video-text pairs

    2. Vision encoderFrozen

      NFNet-F6, contrastively pre-trained first

    3. Perceiver ResamplerTrainable

      64 visual tokens per image

    4. Gated cross-attentionTrainable

      new layers, tanh(α) gate with α = 0 at start

      LM blocksFrozen

      Chinchilla, interleaved with the new layers

    5. Training signal

      next-token loss on text after each image

    Flamingo paper walkthrough
  • Trainableweights update in this stage
  • Frozenruns, but its weights stay fixed
  • Not usednot part of this stage

The three bridge models all start from a vision encoder trained with a contrastive image-text loss: BLIP-2 and LLaVA reuse CLIP-family ViTs, and Flamingo pre-trains its own NFNet that way before freezing it.

In BLIP-2's first stage the LLM is not in the loop at all. The Q-Former learns to extract text-relevant features against its own text transformer, so by the time it meets the LLM it already produces visual features aligned with language. LLaVA reaches a similar starting point by training only its projector on captions before it lets the LLM change.

Side by side

CLIP, BLIP-2, LLaVA and Flamingo compared
Training objectiveCLIPSymmetric contrastive (InfoNCE) loss over the N × N image-text similarities in a batch.BLIP-2Stage 1: image-text contrastive, image-text matching and image-grounded generation. Stage 2: language modeling through the LLM.LLaVANext-token prediction in both stages: captions in stage 1, instruction answers in stage 2.FlamingoNext-token prediction on text, conditioned on the images that precede it.
Alignment mechanismCLIPTwo encoders with linear projections into one shared embedding space.BLIP-2Q-Former: 32 learned queries cross-attend to image features; a linear layer maps them into the LLM input as soft visual prompts.LLaVAA projector (linear, or a two-layer MLP in LLaVA-1.5) maps every patch feature into the word-embedding space.FlamingoPerceiver Resampler, read through tanh-gated cross-attention layers between the LM blocks.
Frozen and trainableCLIPNothing frozen: both encoders and both projections train from scratch.BLIP-2Image encoder and LLM frozen. The Q-Former (188M parameters) trains in both stages, the projection in stage 2.LLaVAVision encoder always frozen. Stage 1 trains the projector; stage 2 trains the projector and the LLM.FlamingoVision encoder and LM blocks frozen. The resampler and the gated cross-attention layers train.
Visual tokens seen by the language sideCLIPNone: one embedding vector per image.BLIP-232 per image.LLaVA256 at 224 px; 576 at 336 px in LLaVA-1.5.Flamingo64 per image, through cross-attention rather than the input sequence.
DataCLIP400M web image-text pairs, in batches of 32,768.BLIP-2129M images with web or synthetic captions, used in both stages.LLaVA595K CC3M caption pairs, then 158K GPT-4-generated instruction samples. LLaVA-1.5: 558K and 665K.Flamingo43M interleaved web pages, 1.8B plus 312M image-text pairs, and 27M video-text pairs.
Downstream fitCLIPRetrieval, zero-shot classification, and a vision encoder for other models. Does not generate text.BLIP-2Captioning and visual question answering with a frozen LLM; the stage-1 model, without the LLM, handles image-text retrieval.LLaVAVisual chat and instruction following over a single image.FlamingoFew-shot captioning, question answering and dialogue over interleaved images and video.

How each model aligns

CLIP: contrastive dual encoders

For a batch of N image-text pairs, CLIP computes the N × N matrix of cosine similarities between image and text embeddings, scales it by a learned temperature, and applies cross-entropy along the rows and along the columns, with the matching pair as the target each time. The symmetric loss, a form of InfoNCE (contrastive loss), treats the other N − 1 captions in the batch as negatives for each image, and the other images as negatives for each caption. Both encoders train from scratch; neither starts from ImageNet or language-model weights. Zero-shot classification follows directly: embed "a photo of a {label}" for each class and pick the class whose text embedding is closest to the image embedding. Deep dive: CLIP paper walkthrough.

BLIP-2: a Q-Former between frozen models

The Q-Former holds 32 learned query vectors that cross-attend to the frozen image encoder's features, so it emits 32 outputs whatever the image. Stage 1 trains it on three objectives at once, image-text contrastive (ITC), image-text matching (ITM) and image-grounded text generation (ITG), each with its own attention mask between queries and text. Stage 2 attaches the frozen LLM: a fully connected layer maps the 32 query outputs to the LLM's embedding size, and they are prepended to the text as soft visual prompts. The Q-Former and the projection train under the LLM's language-modeling loss, and the LLM does not change. Because stage 1 involves no LLM, the paper pairs the same approach with decoder-only OPT models and encoder-decoder FlanT5 models. Deep dive: BLIP-2 paper walkthrough.

LLaVA: patches projected into the word-embedding space

LLaVA maps each patch feature of a frozen CLIP ViT-L/14 through a trainable projection into the word-embedding space of a Vicuna LLM, and places these visual tokens in the input sequence alongside the text tokens. The original uses one linear layer; LLaVA-1.5 (Liu et al., 2023) uses a two-layer MLP and a 336-pixel encoder. Stage 1 trains only the projector on about 595K image-caption pairs filtered from CC3M, with the vision encoder and the LLM frozen. Stage 2 keeps the vision encoder frozen and fine-tunes the projector and the LLM on 158K instruction-following samples, which text-only GPT-4 generated from COCO captions and bounding boxes. The design bets that the LLM can do the cross-modal reasoning once the projector makes the patch features legible to it. Deep dive: Visual Instruction Tuning (LLaVA) walkthrough.

Flamingo: gated cross-attention into a frozen LM

Flamingo's Perceiver Resampler turns a variable number of vision features into 64 visual tokens per image or video. New GATED XATTN-DENSE layers, trained from scratch, sit between the frozen LM blocks: text tokens cross-attend to the visual tokens, and each new layer's output is scaled by tanh(α) before it joins the residual stream. With α initialized to 0, the model starts out computing exactly what the frozen LM computes and admits visual information only as training opens the gates. A mask lets each text token attend only to the image that most recently preceded it, so one sequence can interleave many images and texts. Training on 43M interleaved web pages, alongside image-text and video-text pairs, gives Flamingo few-shot in-context learning: a prompt with a few image, question and answer examples defines a new task without any weight update. Deep dive: Flamingo paper walkthrough.

Choosing an approach

  • Retrieval, search, or zero-shot classification over a label set: a joint embedding model such as CLIP. Embeddings are cheap to index and compare, and nothing needs to generate text.
  • Captioning and visual question answering on a small training budget, with the LLM's text behavior intact: a frozen-LLM bridge such as BLIP-2. Only the Q-Former and the projection train.
  • Visual chat and instruction following: the LLaVA recipe of a projector plus a fine-tuned LLM. It has the simplest connector, and it adapts the LLM itself to multimodal instructions.
  • Many images or video frames interleaved with text, and new tasks from a few examples: cross-attention into a frozen LM, as in Flamingo. Visual tokens do not occupy the LM's input sequence, and the zero-initialized gates protect the frozen LM at the start of training.

Primary sources

If you found this explanation helpful, consider sharing it with others.

Mastodon