LIVE
TutorMoments: Do AI tutors know when to help and when to hold back?07/08/26 · Hugging Face|Responding to the next frontier of critical cyber capabilities07/08/26 · OpenAI|How HSP GRUPPE builds AI capabilities for tax advisory07/08/26 · OpenAI|Improving Fable 5's biology safeguards07/08/26 · Anthropic|AMD acquires Taalas to boost inference performance by etching models in silicon06/08/26|xAI, SpaceX, and the Race for AI Buildout06/08/26 · xAI|The Bitter Lesson of Tool Calling06/08/26 · OpenAI|The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping06/08/26 · Google|Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users06/08/26 · OpenAI|Software development with AI is starting to feel like cooking steak06/08/26|WeatherNext: AI model achieves breakthrough in forecasting cyclones06/08/26 · Google|Into the Omniverse: How Open World Models Push the Frontier of Physical AI06/08/26 · NVIDIA|TutorMoments: Do AI tutors know when to help and when to hold back?07/08/26 · Hugging Face|Responding to the next frontier of critical cyber capabilities07/08/26 · OpenAI|How HSP GRUPPE builds AI capabilities for tax advisory07/08/26 · OpenAI|Improving Fable 5's biology safeguards07/08/26 · Anthropic|AMD acquires Taalas to boost inference performance by etching models in silicon06/08/26|xAI, SpaceX, and the Race for AI Buildout06/08/26 · xAI|The Bitter Lesson of Tool Calling06/08/26 · OpenAI|The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping06/08/26 · Google|Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users06/08/26 · OpenAI|Software development with AI is starting to feel like cooking steak06/08/26|WeatherNext: AI model achieves breakthrough in forecasting cyclones06/08/26 · Google|Into the Omniverse: How Open World Models Push the Frontier of Physical AI06/08/26 · NVIDIA|
ResearchGoogle

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only…

August 6, 20261 min readPublished byarXiv

Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While existing programmatic benchmarks offer better control, they score only the final answer rather than auditing reported events against executable ground truth. To bridge this gap, we introduce trace-grounded parametric profiling for event counting in three controlled video tasks: bouncing-ball wall contacts, visual blinks, and categorical state transitions. Across 2,190 videos, we vary event count N and frequency F while holding rendering fixed. Each video includes an executable event trace for capability-surface estimation and timestamp-level evaluation. Our results reveal a staged temporal failure. At an 80% reliability threshold, Gemini 3.6 Flash reliably counts persistent state transitions up to 12 events at 0.5 and 1.0 Hz, yet demonstrates no reliable positive-count region for transient blinking events. Thus, event representation dictates whether a model initially accesses evidence -- a limitation that compounds as count and frequency increase. In the high-count, high-frequency regime, only 0.2% of final counts are correct and the model recovers just 18.1% of true events. To test if visual access is the primary bottleneck, we increase sampling rate. Although this boosts Bounce Ball accuracy from 19.6% to 29.3%, the reported sequence agrees with ground truth only 3.7% of the time. Extra frames can therefore inflate final scores without producing faithful event recovery. Different prompting strategies yield similarly limited gains, and real-world video evaluations show the same concentration of success at low event counts. Ultimately, trace-grounded profiling shifts video evaluation from aggregate accuracy metrics to a detailed diagnostic of where temporal reasoning fails.

Tags
reasoningbenchmark

Read also

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping · nAIvigate