Skip to main content

Vision-Language Models

Explore machine learning papers and reviews related to Vision-Language Models. Find insights, analysis, and implementation details.

  • Tagged with
  • 8 papers
Back to all papers

Papers Related to Vision-Language Models

arXiv 2024

Qwen2-VL: Vision-Language Perception at Any Resolution

Peng Wang, Shuai Bai, +17

How Qwen2-VL perceives images and video at any resolution with naive dynamic resolution (variable visual tokens) and M-RoPE, a multimodal rotary position embedding that decomposes position into temporal, height, and width components.

Mastodon