Independent research
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
WeaveBench is a long-horizon, hybrid-interface benchmark for computer-use agents, comprising 114 tasks across 8 real-world work domains. Each task requires agents to combine GUI observations/actions with CLI/code operations within a single trajectory, satisfying three admission criteria: channel non-substitutability, long-horizon execution, and…
Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu, et al.- Published
- Jun 2026
- Citations
- 2
- Code
- 160 stars
