Skip to content
Daily PaperJul 13, 2026

← Archive

SOTA with your eyes closed?

Video-Oasis: Rethinking Evaluation of Video Understanding

SOTA with your eyes closed? — 1 / 7
01
SOTA with your eyes closed? — 2 / 7
02
SOTA with your eyes closed? — 3 / 7
03
SOTA with your eyes closed? — 4 / 7
04
SOTA with your eyes closed? — 5 / 7
05
SOTA with your eyes closed? — 6 / 7
06
SOTA with your eyes closed? — 7 / 7
07
01/07

Do the high scores of today’s video AI models really come from understanding video?

This paper argues that 55% of existing video benchmarks can be solved without the video or any temporal context at all.

Three questions put to the benchmarks

Video-Oasis is a diagnostic suite aimed at the benchmarks themselves.

  • (a) Can it be solved with the video removed? — Visual Dependency Test
  • (b) Can it be solved with the frames shuffled? — Temporal Dependency Test
  • (c) Are the questions and answers free of errors and ambiguity? — Ambiguity Verification

What they found

Once the shortcuts were removed, model performance dropped sharply — exposing how hard genuine video understanding still is.

Tags

  • #ai-papers
  • #paper-review
  • #multimodal-ai
  • #computer-vision
View on Instagram