LIVE
Disrupting a Criminal Scam Operation04/08/26 · OpenAI|The Download: reward hacking explained, and suspected Iranian cyberattacks03/08/26 · OpenAI|Here’s why AI agents lie and cheat to reach their goals03/08/26 · OpenAI|OpenAI's super PAC is funding AI-generated news site attacking industry critics03/08/26 · OpenAI|Show HN: Bor – Open-source policy management for Linux desktops02/08/26 · Microsoft|Anionex/codex-deepseek-vision: 让纯文本模型在 Codex 中无障碍调用内置看图工具(view_image)的方案,附为纯文本 LLM 设计的视觉工具包 | Let text-only models call Codex's built-in view_image seamlessly, plus a vision toolkit designed for text-only LLMs.01/08/26 · DeepSeek|Ten advances in mathematics and theoretical computer science01/08/26 · OpenAI|Advancing responsible AI across Europe31/07/26 · OpenAI|Building abundant intelligence31/07/26 · OpenAI|Disrupting a Criminal Scam Operation04/08/26 · OpenAI|The Download: reward hacking explained, and suspected Iranian cyberattacks03/08/26 · OpenAI|Here’s why AI agents lie and cheat to reach their goals03/08/26 · OpenAI|OpenAI's super PAC is funding AI-generated news site attacking industry critics03/08/26 · OpenAI|Show HN: Bor – Open-source policy management for Linux desktops02/08/26 · Microsoft|Anionex/codex-deepseek-vision: 让纯文本模型在 Codex 中无障碍调用内置看图工具(view_image)的方案,附为纯文本 LLM 设计的视觉工具包 | Let text-only models call Codex's built-in view_image seamlessly, plus a vision toolkit designed for text-only LLMs.01/08/26 · DeepSeek|Ten advances in mathematics and theoretical computer science01/08/26 · OpenAI|Advancing responsible AI across Europe31/07/26 · OpenAI|Building abundant intelligence31/07/26 · OpenAI|
ResearchMicrosoft

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and…

July 30, 20261 min readPublished byarXiv

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.

Tags
multimodalagentsreasoningcodingsafety