EN DIRECT
Third-party cyber evaluations involving OpenAI models04/08/26 · OpenAI|Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent04/08/26 · OpenAI|Mistral's Shieldstral: 3B open-weights model for multimodal moderation04/08/26 · Mistral AI|NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US04/08/26 · NVIDIA|Apple says more ex-employees may have taken confidential data to OpenAI04/08/26 · OpenAI|NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use04/08/26 · NVIDIA|As AI Increases Demands on Memory, Storage Steps Up04/08/26 · NVIDIA|Deploy local agents everywhere with LFM2.5-2.6B04/08/26 · Hugging Face|AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency04/08/26 · NVIDIA|AI-Generated Images Discourage Me from Reading Your Blog04/08/26|Disrupting a Criminal Scam Operation04/08/26 · OpenAI|Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer04/08/26 · Anthropic|Third-party cyber evaluations involving OpenAI models04/08/26 · OpenAI|Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent04/08/26 · OpenAI|Mistral's Shieldstral: 3B open-weights model for multimodal moderation04/08/26 · Mistral AI|NVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US04/08/26 · NVIDIA|Apple says more ex-employees may have taken confidential data to OpenAI04/08/26 · OpenAI|NVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use04/08/26 · NVIDIA|As AI Increases Demands on Memory, Storage Steps Up04/08/26 · NVIDIA|Deploy local agents everywhere with LFM2.5-2.6B04/08/26 · Hugging Face|AI Leaders Propose SAFE Guidelines for Cybersecurity Transparency04/08/26 · NVIDIA|AI-Generated Images Discourage Me from Reading Your Blog04/08/26|Disrupting a Criminal Scam Operation04/08/26 · OpenAI|Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer04/08/26 · Anthropic|
RechercheOpenAI

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two…

4 août 20261 min de lecturePublié pararXiv

We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.

Tags
multimodalagentsragcodingbenchmark

À lire aussi

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent · nAIvigate