DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration

June 2026 Wenkai Wang*, Tao Xiong*, Jingchen Ni*, Yunpeng Bao*, Xiyun Li, Tianqi Liu, Hongcan Guo, Zilong Huang, Shengyu Zhang Under review at EMNLP 2026 (THU-A)
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration

Overview

As a co-first author, I helped build DeskCraft, a desktop GUI agent benchmark targeting long-horizon professional creative and engineering workflows and proactive human-agent collaboration. It comprises 538 executable tasks across 11 applications with a three-level difficulty taxonomy and a composable human-in-the-loop interaction protocol. Evaluations of 18 agents reveal persistent failures in long-horizon delivery and proactive clarification.