ICML 2022
BLIP: Bootstrapping Language-Image Pre-training
Junnan Li, Dongxu Li, +2
How BLIP unifies vision-language understanding and generation in one model, and bootstraps noisy web data with CapFilt — a captioner that synthesizes captions and a filter that removes noisy image-text pairs.
