arXiv 2024
Qwen2-VL: Vision-Language Perception at Any Resolution
Peng Wang, Shuai Bai, +17
How Qwen2-VL perceives images and video at any resolution with naive dynamic resolution (variable visual tokens) and M-RoPE, a multimodal rotary position embedding that decomposes position into temporal, height, and width components.
