Every vision-language model has to bring images and text to a place where they can be compared or combined. CLIP does it by training two encoders into one shared embedding space. BLIP-2, LLaVA and Flamingo start instead from a pretrained vision encoder and a pretrained language model, and train a bridge that lets the language model read visual features. This page compares the four mechanisms, what each one trains and freezes, and which tasks each one suits. The alignment problem page covers the objective in general, and each model below links to its own paper walkthrough.
TL;DR
- All four models align visual and textual representations. They differ in where the alignment lives and in which weights learn it.
- CLIP (Radford et al., 2021) trains an image encoder and a text encoder from scratch with a symmetric contrastive loss on 400M image-text pairs. Matching images and captions land close together in one embedding space, which gives retrieval and zero-shot classification but no text generation.
- BLIP-2 (Li et al., 2023) keeps a frozen image encoder and a frozen LLM and trains a lightweight Querying Transformer (Q-Former) between them, whose 32 learned queries pull visual features out of the encoder. Stage 1 aligns the Q-Former with text; stage 2 connects it to the LLM through a linear projection.
- LLaVA (Liu et al., 2023) projects the patch features of a frozen CLIP ViT-L/14 into the LLM's word-embedding space, with a linear layer in the original and a two-layer MLP in LLaVA-1.5. Stage 1 trains only the projector; stage 2 also fine-tunes the LLM.
- Flamingo (Alayrac et al., 2022) keeps a frozen vision encoder and a frozen language model and trains a Perceiver Resampler plus new tanh-gated cross-attention layers between the LM blocks. It reads interleaved images and text and learns new tasks from a few examples in its prompt.
The fundamental design choice
Train a joint embedding space
Two encoders, one per modality, each end in a projection to the same vector space. Training pulls matched image-caption pairs together and pushes mismatched pairs apart, so the alignment is geometric: the similarity of an image and a text is the dot product of two unit vectors.
Images and texts are encoded independently, so embeddings can be computed once and indexed, and a classifier is simply a set of text prompts. The cost is that each image and each caption becomes one vector, compared by one score. Nothing in this setup writes text, and a single vector drops detail that a question about the image might need. Contrastive training also leaves the two modalities in separate regions of the shared space, the modality gap.
Project a vision encoder into a language model
The second family keeps a pretrained LLM as the part that reasons and writes, and turns the image into something it can read. Alignment here means the language model can use the visual features. The bridge is trained with the language model's own next-token loss, so it is judged by whether the text that follows the image becomes more likely. This gives captioning, question answering and dialogue, but every query runs the language model, so the result is not a retrieval index.
Three further choices
Within the projection family, three choices separate BLIP-2, LLaVA and Flamingo.
- Where visual information enters. BLIP-2 and LLaVA place visual tokens in the LLM's input sequence, where its existing self-attention reads them like text, with no new layers. Flamingo keeps the input sequence as text and adds new cross-attention layers inside the LM, through which text tokens read the visual tokens.
- Whether the LLM learns. BLIP-2 and Flamingo keep the LLM frozen, which preserves its text abilities and keeps the trainable part small; BLIP-2 reports 54 times fewer trainable parameters than Flamingo-80B. LLaVA fine-tunes the LLM in its second stage, giving up that protection for stronger instruction following. Parameter-efficient methods such as LoRA sit between the two (vision-language adapters).
- How many visual tokens. LLaVA passes every patch through: 256 tokens for a 224-pixel image with ViT-L/14, and 576 at 336 pixels in LLaVA-1.5. BLIP-2 compresses any image to 32 query outputs, and Flamingo's resampler to 64 tokens per image. Fewer tokens use less compute and context; more tokens keep more spatial detail.
The two families are stages of one pipeline more than rival lineages. All three bridge models start from a vision encoder trained with a contrastive image-text loss: BLIP-2 uses CLIP's ViT-L/14 or EVA-CLIP's ViT-g/14, LLaVA uses CLIP's ViT-L/14, and Flamingo pre-trains its own NFNet-F6 contrastively before freezing it.
Four architectures, stage by stage
Each panel follows one model from its input to its training signal. Pick a stage to see which parts learn in it; every part is marked trainable, frozen, or not used, in text and with an icon.
Stage 1: BLIP-2 trains only its Q-Former, with no LLM in the loop, and LLaVA trains only its projector. CLIP and Flamingo train in one stage, so they read the same in both.
CLIP, one stage: trains the image encoder, text encoder and linear projections. BLIP-2, stage 1: trains the Q-Former; frozen: image encoder; not used: linear projection and LLM. LLaVA, stage 1: trains the projector; frozen: vision encoder and LLM. Flamingo, one stage: trains the Perceiver Resampler and gated cross-attention; frozen: vision encoder and LM blocks.
CLIP Radford et al., 2021
Contrastive dual encoder
One stage: contrastive pre-training
- Image
from a web image-text pair
Captionits paired text
- Image encoderTrainable
ResNet or ViT, from scratch
Text encoderTrainableTransformer, from scratch
- Linear projectionsTrainable
both towers into one shared embedding space
- Symmetric contrastive loss
over the N × N image-text similarities; matched pairs pulled together
BLIP-2 Li et al., 2023
Q-Former between two frozen models
Stage 1: representation learning
- Image
with a caption
Captionread by the Q-Former
- Image encoderFrozen
ViT-L/14 (CLIP) or ViT-g/14 (EVA-CLIP)
- Q-FormerTrainable
32 learned queries cross-attend to image features
- Linear projectionNot used
32 query outputs into the LLM input space
- LLMNot used
OPT or FlanT5
- Training signal
ITC, ITM and ITG losses on the Q-Former
LLaVA Liu et al., 2023
Projection into the LLM input
Stage 1: feature alignment
- Image
one per conversation
Textcaption as a brief description
- Vision encoderFrozen
CLIP ViT-L/14 patch features
- ProjectorTrainable
linear (LLaVA) or two-layer MLP (LLaVA-1.5)
- LLMFrozen
Vicuna reads visual and text tokens together
- Training signal
next-token loss on 595K CC3M captions
Flamingo Alayrac et al., 2022
Gated cross-attention into a frozen LM
One stage: training on a multimodal web mixture
- Interleaved images and text
web pages, image-text and video-text pairs
- Vision encoderFrozen
NFNet-F6, contrastively pre-trained first
- Perceiver ResamplerTrainable
64 visual tokens per image
- Gated cross-attentionTrainable
new layers, tanh(α) gate with α = 0 at start
LM blocksFrozenChinchilla, interleaved with the new layers
- Training signal
next-token loss on text after each image
- Trainableweights update in this stage
- Frozenruns, but its weights stay fixed
- Not usednot part of this stage
The three bridge models all start from a vision encoder trained with a contrastive image-text loss: BLIP-2 and LLaVA reuse CLIP-family ViTs, and Flamingo pre-trains its own NFNet that way before freezing it.
In BLIP-2's first stage the LLM is not in the loop at all. The Q-Former learns to extract text-relevant features against its own text transformer, so by the time it meets the LLM it already produces visual features aligned with language. LLaVA reaches a similar starting point by training only its projector on captions before it lets the LLM change.
Side by side
| Aspect | CLIP | BLIP-2 | LLaVA | Flamingo |
|---|---|---|---|---|
| Training objective | CLIPSymmetric contrastive (InfoNCE) loss over the N × N image-text similarities in a batch. | BLIP-2Stage 1: image-text contrastive, image-text matching and image-grounded generation. Stage 2: language modeling through the LLM. | LLaVANext-token prediction in both stages: captions in stage 1, instruction answers in stage 2. | FlamingoNext-token prediction on text, conditioned on the images that precede it. |
| Alignment mechanism | CLIPTwo encoders with linear projections into one shared embedding space. | BLIP-2Q-Former: 32 learned queries cross-attend to image features; a linear layer maps them into the LLM input as soft visual prompts. | LLaVAA projector (linear, or a two-layer MLP in LLaVA-1.5) maps every patch feature into the word-embedding space. | FlamingoPerceiver Resampler, read through tanh-gated cross-attention layers between the LM blocks. |
| Frozen and trainable | CLIPNothing frozen: both encoders and both projections train from scratch. | BLIP-2Image encoder and LLM frozen. The Q-Former (188M parameters) trains in both stages, the projection in stage 2. | LLaVAVision encoder always frozen. Stage 1 trains the projector; stage 2 trains the projector and the LLM. | FlamingoVision encoder and LM blocks frozen. The resampler and the gated cross-attention layers train. |
| Visual tokens seen by the language side | CLIPNone: one embedding vector per image. | BLIP-232 per image. | LLaVA256 at 224 px; 576 at 336 px in LLaVA-1.5. | Flamingo64 per image, through cross-attention rather than the input sequence. |
| Data | CLIP400M web image-text pairs, in batches of 32,768. | BLIP-2129M images with web or synthetic captions, used in both stages. | LLaVA595K CC3M caption pairs, then 158K GPT-4-generated instruction samples. LLaVA-1.5: 558K and 665K. | Flamingo43M interleaved web pages, 1.8B plus 312M image-text pairs, and 27M video-text pairs. |
| Downstream fit | CLIPRetrieval, zero-shot classification, and a vision encoder for other models. Does not generate text. | BLIP-2Captioning and visual question answering with a frozen LLM; the stage-1 model, without the LLM, handles image-text retrieval. | LLaVAVisual chat and instruction following over a single image. | FlamingoFew-shot captioning, question answering and dialogue over interleaved images and video. |
How each model aligns
CLIP: contrastive dual encoders
For a batch of N image-text pairs, CLIP computes the N × N matrix of cosine similarities between image and text embeddings, scales it by a learned temperature, and applies cross-entropy along the rows and along the columns, with the matching pair as the target each time. The symmetric loss, a form of InfoNCE (contrastive loss), treats the other N − 1 captions in the batch as negatives for each image, and the other images as negatives for each caption. Both encoders train from scratch; neither starts from ImageNet or language-model weights. Zero-shot classification follows directly: embed "a photo of a {label}" for each class and pick the class whose text embedding is closest to the image embedding. Deep dive: CLIP paper walkthrough.
BLIP-2: a Q-Former between frozen models
The Q-Former holds 32 learned query vectors that cross-attend to the frozen image encoder's features, so it emits 32 outputs whatever the image. Stage 1 trains it on three objectives at once, image-text contrastive (ITC), image-text matching (ITM) and image-grounded text generation (ITG), each with its own attention mask between queries and text. Stage 2 attaches the frozen LLM: a fully connected layer maps the 32 query outputs to the LLM's embedding size, and they are prepended to the text as soft visual prompts. The Q-Former and the projection train under the LLM's language-modeling loss, and the LLM does not change. Because stage 1 involves no LLM, the paper pairs the same approach with decoder-only OPT models and encoder-decoder FlanT5 models. Deep dive: BLIP-2 paper walkthrough.
LLaVA: patches projected into the word-embedding space
LLaVA maps each patch feature of a frozen CLIP ViT-L/14 through a trainable projection into the word-embedding space of a Vicuna LLM, and places these visual tokens in the input sequence alongside the text tokens. The original uses one linear layer; LLaVA-1.5 (Liu et al., 2023) uses a two-layer MLP and a 336-pixel encoder. Stage 1 trains only the projector on about 595K image-caption pairs filtered from CC3M, with the vision encoder and the LLM frozen. Stage 2 keeps the vision encoder frozen and fine-tunes the projector and the LLM on 158K instruction-following samples, which text-only GPT-4 generated from COCO captions and bounding boxes. The design bets that the LLM can do the cross-modal reasoning once the projector makes the patch features legible to it. Deep dive: Visual Instruction Tuning (LLaVA) walkthrough.
Flamingo: gated cross-attention into a frozen LM
Flamingo's Perceiver Resampler turns a variable number of vision features into 64 visual tokens per image or video. New GATED XATTN-DENSE layers, trained from scratch, sit between the frozen LM blocks: text tokens cross-attend to the visual tokens, and each new layer's output is scaled by tanh(α) before it joins the residual stream. With α initialized to 0, the model starts out computing exactly what the frozen LM computes and admits visual information only as training opens the gates. A mask lets each text token attend only to the image that most recently preceded it, so one sequence can interleave many images and texts. Training on 43M interleaved web pages, alongside image-text and video-text pairs, gives Flamingo few-shot in-context learning: a prompt with a few image, question and answer examples defines a new task without any weight update. Deep dive: Flamingo paper walkthrough.
Choosing an approach
- Retrieval, search, or zero-shot classification over a label set: a joint embedding model such as CLIP. Embeddings are cheap to index and compare, and nothing needs to generate text.
- Captioning and visual question answering on a small training budget, with the LLM's text behavior intact: a frozen-LLM bridge such as BLIP-2. Only the Q-Former and the projection train.
- Visual chat and instruction following: the LLaVA recipe of a projector plus a fine-tuned LLM. It has the simplest connector, and it adapts the LLM itself to multimodal instructions.
- Many images or video frames interleaved with text, and new tasks from a few examples: cross-attention into a frozen LM, as in Flamingo. Visual tokens do not occupy the LM's input sequence, and the zero-initialized gates protect the frozen LM at the start of training.
Primary sources
- Learning Transferable Visual Models From Natural Language Supervision (Radford et al., 2021), CLIP
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (Li et al., 2023)
- Visual Instruction Tuning (Liu et al., 2023), LLaVA
- Improved Baselines with Visual Instruction Tuning (Liu et al., 2023), LLaVA-1.5
- Flamingo: a Visual Language Model for Few-Shot Learning (Alayrac et al., 2022)
Related concepts
How vision-language models align visual and text representations using contrastive learning, cross-modal attention, and CLIP-style training.
The modality gap in CLIP and vision-language models: why image and text embeddings occupy separate regions despite contrastive training.
Discover how multimodal vision-language models like CLIP, ALIGN, and LLaVA scale with data, parameters, and compute following Chinchilla-style power laws.
Master LoRA, bottleneck adapters, and prefix tuning for parameter-efficient fine-tuning of vision-language models like LLaVA with minimal compute and memory.
Interactive visualization of LLM context windows - sliding windows, expanding contexts, and attention patterns that define model memory limits.
Understand cross-attention, the mechanism that enables transformers to align and fuse information from different sources, sequences, or modalities.
