Skip to main content

Video Understanding

Explore machine learning papers and reviews related to Video Understanding. Find insights, analysis, and implementation details.

  • Tagged with
  • 2 papers
Back to all papers

Papers Related to Video Understanding

arXiv 2025

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Mido Assran, Adrien Bardes, +27

How V-JEPA 2 scales self-supervised video learning to 1M+ hours with mask denoising and 3D-RoPE, then extends to V-JEPA 2-AC — an action-conditioned world model that enables zero-shot robotic planning from just 62 hours of unlabeled video.

TMLR 2024

V-JEPA: Learning Video Representations by Predicting in Latent Space

Adrien Bardes, Quentin Garrido, +6

How V-JEPA learns powerful video representations by predicting masked spatiotemporal regions in embedding space rather than reconstructing pixels, achieving state-of-the-art frozen features with superior label efficiency.

Mastodon