Motion-o: Trajectory-Grounded Video Reasoning

📰 ArXiv cs.AI

arXiv:2603.18856v2 Announce Type: replace-cross Abstract: Recent video reasoning models increasingly produce spatio-temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grounding \emph{where} and \emph{when} evidence appears, they often leave the motion connecting observations, the \textit{how}, implicit. This makes dynamic and trajectory-dependent claims difficult to supervise, verify, or penalize when unsupported by the video. We

Published 11 May 2026

Full Article

Title: Motion-o: Trajectory-Grounded Video Reasoning

Abstract:
arXiv:2603.18856v2 Announce Type: replace-cross Abstract: Recent video reasoning models increasingly produce spatio-temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grounding \emph{where} and \emph{when} evidence appears, they often leave the motion connecting observations, the \textit{how}, implicit. This makes dynamic and trajectory-dependent claims difficult to supervise, verify, or penalize when unsupported by the video. We
Read full paper → ← Back to Reads