Daily AI Catchup
H-CompanyComputer-UseBenchmarksAgents

H Company's Holo3.1 Now Beats Both OpenAI and Anthropic on Desktop Computer-Use Tasks

Holo3.1 scores 78.85% on the OSWorld desktop-task benchmark, beating both OpenAI and Anthropic's computer-use offerings. Desktop GUI agents are rapidly becoming a real, competitive capability category rather than a research curiosity.

Why it matters

๐Ÿ’ป Developer ยท If you've been using OpenAI or Anthropic's computer-use APIs for desktop automation, Holo3.1 is now worth a direct head-to-head benchmark on your actual use case.

๐Ÿ“ฆ Product ยท An independent, open challenger beating both major labs on a real benchmark suggests desktop automation is maturing into a competitive market rather than a single-vendor category โ€” worth revisiting your assumptions.

๐ŸŽจ Design ยท No direct design impact โ€” this is a backend automation capability.

๐Ÿ“ˆ Business ยท A smaller, independent company outperforming both OpenAI and Anthropic on a specific capability is a reminder that frontier-lab dominance doesn't automatically extend to every specialized task.

๐Ÿค” Just Curious ยท A lesser-known AI company just built a computer-controlling AI that's now better at navigating desktop apps and menus than the equivalent tools from both OpenAI and Anthropic.

Sources: H Company's Holo3.1 beats OpenAI and Anthropic on desktop tasks