OSWorld 2.0 Shows AI Agents Still Break on Long Tasks
Computer Use Agents A harder test showed that AI agents controlling a computer still fall apart on tasks longer than a few dozen steps. By Shashi Bellamkonda · August 22, 2026 20.6% the top AI model's score on OSWorld 2.0's harder, longer task…