Daily PaperJul 13, 2026
01/07
Do the high scores of today’s video AI models really come from understanding video?
This paper argues that 55% of existing video benchmarks can be solved without the video or any temporal context at all.
Three questions put to the benchmarks
Video-Oasis is a diagnostic suite aimed at the benchmarks themselves.
- (a) Can it be solved with the video removed? — Visual Dependency Test
- (b) Can it be solved with the frames shuffled? — Temporal Dependency Test
- (c) Are the questions and answers free of errors and ambiguity? — Ambiguity Verification
What they found
Once the shortcuts were removed, model performance dropped sharply — exposing how hard genuine video understanding still is.
Tags
- #ai-papers
- #paper-review
- #multimodal-ai
- #computer-vision






