arXiv:2606.09833cs.HCcs.AI2026-06

构建真实人机协作评估框架,揭示协同技能核心影响因素。

CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks

论文配图:CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks
图 1 · 摘自论文原文
  • 通过真实工作者与AI协作收集数据,量化双方贡献
  • 实证显示实用经验是协作能力主因,协作提升用户AI素养
  • 适合关注人机协同效能与智能助手设计的研究者

AI代理正在重塑工作方式,但人机协作仍缺乏真实任务评价体系,受限于真实人类数据获取难与个体差异。我们提出CollabSkill框架,将真实从业者与AI代理配对于其职业背景匹配的任务中,收集涵盖经济价值任务复杂性及真实使用模式的数据。为应对个体差异,采用贝叶斯技能评分系统分离并量化人类与AI的技能贡献。基于93名工作者完成的386次会话、超过1500个提示的分析发现:在协作评估中,Claude Code排名领先,而传统自主基准中Codex居首;人类方面,实践经验是协作技能的核心驱动力,实际协作显著提升工人对AI的认知水平。本研究旨在推动社区建立系统化的人机协作评估机制,助力开发真正增强人类能力的AI代理。

原文摘要 · Abstract (English)

AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent collaboration both in preserving human agency and generating economic value, this paradigm remains largely absent from occupational task evaluation, hindered by the difficulty of gathering real human data and accounting for inter-human variability. We introduce CollabSkill, a framework for evaluating human-agent collaboration on real-world occupational tasks. CollabSkill pairs real human workers with AI agents on tasks matched to their occupational background, collecting data that capture the complexity of economically valuable tasks and the usage patterns of real workers. To account for inter-human variability, CollabSkill employs a Bayesian skill rating system to disentangle and quantify the skill contributions of both humans and AI agents. Drawing on over 1,500 prompts from 386 working sessions contributed by 93 human workers, our analysis yields insights on two fronts: on the agent side, rankings on CollabSkill diverge meaningfully from those of existing fully autonomous benchmarks where Codex leads, with Claude Code ranking first; on the human side, CollabSkill reveals that practical experience emerges as the primary driver of collaboration skill, with hands-on collaboration meaningfully shifting workers' AI literacy. Together, we hope CollabSkill enables the community to invest in systematic evaluation of human-agent collaboration and spurs development efforts aimed at building AI agents that genuinely augment human workers.

人机协作评估框架真实任务技能量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。