Skip to main content

Multimodal Scaling Laws

Summary
Discover how multimodal vision-language models like CLIP, ALIGN, and LLaVA scale with data, parameters, and compute following Chinchilla-style power laws.

Multimodal Scaling Laws

Multimodal models exhibit unique scaling behaviors that differ from single-modality systems. Understanding these laws is crucial for efficient training and optimal resource allocation.

Interactive Scaling Explorer

The Chinchilla Law for Multimodal

Hoffmann et al. (2022, Chinchilla) model the loss of a model with N parameters trained on D tokens as an irreducible loss plus one power-law term for each:

L(N, D) = E + ANα + BDβ

Where:

  • N = Number of parameters
  • D = Number of training tokens
  • E = Irreducible loss of the data
  • A, B, α, β = Fitted constants

Compute is not a separate term: training costs about C ≈ 6ND FLOPs, so a fixed compute budget sets the trade-off between N and D. Chinchilla's parametric fit (Approach 3, fitted on text only) gives E = 1.69, A = 406.4, B = 410.7, α = 0.34 and β = 0.28. Aghajanyan et al. (2023) fit the same form separately for each modality of a generative token model and find very different constants: for image-text tokens α = 0.12 and β = 0.11, against 0.18 and 0.22 for text in the same study.

Key Scaling Relationships

1. Data Scaling

In Chinchilla's text-only fit, the data term is:

BDβ = 410.7D0.28

Vision-language data scales differently: in Aghajanyan et al.'s generative models the image-text data exponent is 0.11 against 0.22 for text, and Cherti et al. (2023) find that CLIP models trained on OpenAI's WIT and on LAION scale at different rates.

Implications:

  • Halving the data term takes 21/0.28 ≈ 12 times more data
  • Quality matters more than quantity at scale
  • Diverse data sources critical for generalization

2. Model Scaling

Parameters scale with diminishing returns. In the same fit, the parameter term is:

ANα = 406.4N0.34

Key insights:

  • The vision-language adapter and vision encoder are not a fixed share of the model: in Flamingo they take about 1.8B of 3.2B parameters (Flamingo-3B) but about 10.6B of 80B (Flamingo-80B)
  • Flamingo adds a cross-attention block before every language-model layer at 3B, every 4th at 9B and every 7th at 80B; every 4th trained 66% faster than every layer for a 1.9% lower overall score
  • None of the cited papers derives an optimal vision:language parameter ratio; Flamingo keeps its 435M vision encoder fixed from 3B to 80B and grows the language model and cross-attention layers

3. Compute Scaling

FLOPs follow predictable patterns. Kaplan et al. (2020) fit language-model loss against the minimum compute needed to reach it:

L(Cmin) = (CcminCmin)0.050, Ccmin ≈ 3.1 × 108\ \text{PF-days}

For CLIP-style contrastive models, Cherti et al. (2023) fit zero-shot ImageNet error as a power law in training compute, with an exponent of −0.11 for OpenCLIP on LAION and −0.16 for OpenAI CLIP on WIT.

Observations:

  • Compute-optimal training uses about 20 tokens per parameter (from Chinchilla's Table 3: 1B parameters for 20.2B tokens)
  • Parameters and tokens should grow in about equal proportion with compute (Chinchilla's three approaches give exponents of 0.46 to 0.54), where Kaplan et al. had parameters growing as C0.73
  • Vision processing is compute-intensive
  • Batch size affects scaling efficiency

Empirical Findings

Model Comparisons

ModelParametersDataComputePerformance
CLIP-B/32~151M (88M image + 63M text)400M pairsNot reported76.1% ImageNet linear probe
CLIP-L/14~428M (304M image + 124M text)400M pairs256 V100 × 12 days83.9% ImageNet linear probe
ALIGNNot reported (EfficientNet-L2 + BERT-Large)1.8B pairs1024 TPUv3 cores, 1.2M steps85.5% ImageNet linear probe
Flamingo80B1.8B + 312M pairs, 27M videos, 43M webpages1536 TPUv4 × 15 days67.6% VQAv2 test-dev (32-shot)
LLaVA-1.513B1.2M samples8 A100 × ~1 day80.0% VQAv2

Unique Multimodal Phenomena

1. Modality Imbalance

When scaling is imbalanced, the modality gap widens:

  • Vision >> Language: Overfitting on visual features
  • Language >> Vision: Poor grounding, hallucinations
  • Balance: none of the cited papers gives an optimal vision:language:compute ratio. Aghajanyan et al. (2023) instead model how two modalities in one model compete or help each other as model size and data grow; for speech and text their law predicted the competition would end at about 28B parameters and 45B tokens, and a 30B model trained on 50B tokens crossed that point

2. Emergent Abilities

None of the cited papers reports abilities switching on at fixed parameter counts; they report steady gains with scale:

  • CLIP: average zero-shot error follows a smooth log-log trend across a 44× range of model compute, though single datasets are much noisier
  • Flamingo: few-shot performance improves from 3B to 9B to 80B parameters, and the largest model makes better use of extra shots

3. Data Efficiency Paradox

Multimodal models show:

  • Better few-shot learning than unimodal
  • Worse data efficiency during pre-training
  • No cited paper gives a minimum number of pairs; CLIP collected 400M because filtering YFCC100M left only 15M usable pairs, about the size of ImageNet

Practical Guidelines

When to Scale What

Scale Data When:

  • Downstream tasks are diverse
  • Generalization is critical
  • Have compute constraints

Scale Model When:

  • Need complex reasoning
  • Have sufficient data
  • Can afford inference cost

Scale Compute When:

  • Time is critical
  • Have parallel resources
  • Optimizing for convergence

Cost-Performance Trade-offs

StrategyCostPerformanceBest For
Data-heavyLowGoodNarrow domains
Model-heavyHighExcellentGeneral purpose
Compute-heavyMediumGoodRapid iteration
BalancedMediumVery GoodMost use cases

References

  • Kaplan et al. "Scaling Laws for Neural Language Models"
  • Hoffmann et al. "Training Compute-Optimal Large Language Models" (Chinchilla)
  • Cherti et al. "Reproducible Scaling Laws for Contrastive Language-Image Learning"
  • Aghajanyan et al. "Scaling Laws for Generative Mixed-Modal Language Models"
  • Radford et al. "Learning Transferable Visual Models From Natural Language Supervision" (CLIP)
  • Jia et al. "Scaling Up Visual and Vision-Language Representation Learning" (ALIGN)
  • Alayrac et al. "Flamingo: a Visual Language Model for Few-Shot Learning"
  • Liu et al. "Visual Instruction Tuning" (LLaVA)

If you found this explanation helpful, consider sharing it with others.

Mastodon