Multimodal Scaling Laws
Multimodal models exhibit unique scaling behaviors that differ from single-modality systems. Understanding these laws is crucial for efficient training and optimal resource allocation.
Interactive Scaling Explorer
The Chinchilla Law for Multimodal
Hoffmann et al. (2022, Chinchilla) model the loss of a model with N parameters trained on D tokens as an irreducible loss plus one power-law term for each:
Where:
- N = Number of parameters
- D = Number of training tokens
- E = Irreducible loss of the data
- A, B, α, β = Fitted constants
Compute is not a separate term: training costs about C ≈ 6ND FLOPs, so a fixed compute budget sets the trade-off between N and D. Chinchilla's parametric fit (Approach 3, fitted on text only) gives E = 1.69, A = 406.4, B = 410.7, α = 0.34 and β = 0.28. Aghajanyan et al. (2023) fit the same form separately for each modality of a generative token model and find very different constants: for image-text tokens α = 0.12 and β = 0.11, against 0.18 and 0.22 for text in the same study.
Key Scaling Relationships
1. Data Scaling
In Chinchilla's text-only fit, the data term is:
Vision-language data scales differently: in Aghajanyan et al.'s generative models the image-text data exponent is 0.11 against 0.22 for text, and Cherti et al. (2023) find that CLIP models trained on OpenAI's WIT and on LAION scale at different rates.
Implications:
- Halving the data term takes 21/0.28 ≈ 12 times more data
- Quality matters more than quantity at scale
- Diverse data sources critical for generalization
2. Model Scaling
Parameters scale with diminishing returns. In the same fit, the parameter term is:
Key insights:
- The vision-language adapter and vision encoder are not a fixed share of the model: in Flamingo they take about 1.8B of 3.2B parameters (Flamingo-3B) but about 10.6B of 80B (Flamingo-80B)
- Flamingo adds a cross-attention block before every language-model layer at 3B, every 4th at 9B and every 7th at 80B; every 4th trained 66% faster than every layer for a 1.9% lower overall score
- None of the cited papers derives an optimal vision:language parameter ratio; Flamingo keeps its 435M vision encoder fixed from 3B to 80B and grows the language model and cross-attention layers
3. Compute Scaling
FLOPs follow predictable patterns. Kaplan et al. (2020) fit language-model loss against the minimum compute needed to reach it:
For CLIP-style contrastive models, Cherti et al. (2023) fit zero-shot ImageNet error as a power law in training compute, with an exponent of −0.11 for OpenCLIP on LAION and −0.16 for OpenAI CLIP on WIT.
Observations:
- Compute-optimal training uses about 20 tokens per parameter (from Chinchilla's Table 3: 1B parameters for 20.2B tokens)
- Parameters and tokens should grow in about equal proportion with compute (Chinchilla's three approaches give exponents of 0.46 to 0.54), where Kaplan et al. had parameters growing as C0.73
- Vision processing is compute-intensive
- Batch size affects scaling efficiency
Empirical Findings
Model Comparisons
| Model | Parameters | Data | Compute | Performance |
|---|---|---|---|---|
| CLIP-B/32 | ~151M (88M image + 63M text) | 400M pairs | Not reported | 76.1% ImageNet linear probe |
| CLIP-L/14 | ~428M (304M image + 124M text) | 400M pairs | 256 V100 × 12 days | 83.9% ImageNet linear probe |
| ALIGN | Not reported (EfficientNet-L2 + BERT-Large) | 1.8B pairs | 1024 TPUv3 cores, 1.2M steps | 85.5% ImageNet linear probe |
| Flamingo | 80B | 1.8B + 312M pairs, 27M videos, 43M webpages | 1536 TPUv4 × 15 days | 67.6% VQAv2 test-dev (32-shot) |
| LLaVA-1.5 | 13B | 1.2M samples | 8 A100 × ~1 day | 80.0% VQAv2 |
Unique Multimodal Phenomena
1. Modality Imbalance
When scaling is imbalanced, the modality gap widens:
- Vision >> Language: Overfitting on visual features
- Language >> Vision: Poor grounding, hallucinations
- Balance: none of the cited papers gives an optimal vision:language:compute ratio. Aghajanyan et al. (2023) instead model how two modalities in one model compete or help each other as model size and data grow; for speech and text their law predicted the competition would end at about 28B parameters and 45B tokens, and a 30B model trained on 50B tokens crossed that point
2. Emergent Abilities
None of the cited papers reports abilities switching on at fixed parameter counts; they report steady gains with scale:
- CLIP: average zero-shot error follows a smooth log-log trend across a 44× range of model compute, though single datasets are much noisier
- Flamingo: few-shot performance improves from 3B to 9B to 80B parameters, and the largest model makes better use of extra shots
3. Data Efficiency Paradox
Multimodal models show:
- Better few-shot learning than unimodal
- Worse data efficiency during pre-training
- No cited paper gives a minimum number of pairs; CLIP collected 400M because filtering YFCC100M left only 15M usable pairs, about the size of ImageNet
Practical Guidelines
When to Scale What
Scale Data When:
- Downstream tasks are diverse
- Generalization is critical
- Have compute constraints
Scale Model When:
- Need complex reasoning
- Have sufficient data
- Can afford inference cost
Scale Compute When:
- Time is critical
- Have parallel resources
- Optimizing for convergence
Cost-Performance Trade-offs
| Strategy | Cost | Performance | Best For |
|---|---|---|---|
| Data-heavy | Low | Good | Narrow domains |
| Model-heavy | High | Excellent | General purpose |
| Compute-heavy | Medium | Good | Rapid iteration |
| Balanced | Medium | Very Good | Most use cases |
References
- Kaplan et al. "Scaling Laws for Neural Language Models"
- Hoffmann et al. "Training Compute-Optimal Large Language Models" (Chinchilla)
- Cherti et al. "Reproducible Scaling Laws for Contrastive Language-Image Learning"
- Aghajanyan et al. "Scaling Laws for Generative Mixed-Modal Language Models"
- Radford et al. "Learning Transferable Visual Models From Natural Language Supervision" (CLIP)
- Jia et al. "Scaling Up Visual and Vision-Language Representation Learning" (ALIGN)
- Alayrac et al. "Flamingo: a Visual Language Model for Few-Shot Learning"
- Liu et al. "Visual Instruction Tuning" (LLaVA)
Related concepts
How vision-language models align visual and text representations using contrastive learning, cross-modal attention, and CLIP-style training.
The modality gap in CLIP and vision-language models: why image and text embeddings occupy separate regions despite contrastive training.
Master LoRA, bottleneck adapters, and prefix tuning for parameter-efficient fine-tuning of vision-language models like LLaVA with minimal compute and memory.
CLIP, BLIP-2, LLaVA and Flamingo align images with text by different routes. Compare their objectives, bridges, data, and which parts train or stay frozen.
How Flash Attention, Multi-Head Attention (MHA), Grouped-Query Attention (GQA), and Multi-Query Attention (MQA) compare — algorithm vs architecture, KV-cache memory, quality trade-offs, and how to choose for production transformer inference.
Understand cross-attention, the mechanism that enables transformers to align and fuse information from different sources, sequences, or modalities.
