The Vision-Language Alignment Problem
How vision-language models align visual and text representations using contrastive learning, cross-modal attention, and CLIP-style training.
7 min readConcept
Explore machine learning concepts related to CLIP. Clear explanations and practical insights.
How vision-language models align visual and text representations using contrastive learning, cross-modal attention, and CLIP-style training.
CLIP, BLIP-2, LLaVA and Flamingo align images with text by different routes. Compare their objectives, bridges, data, and which parts train or stay frozen.