arXiv:2511.09309cs.HCcs.AI2025-11被引 2

用认知链模型量化界面操作的心理难度,更真实反映任务挑战性。

TaskSense: Cognitive Chain Modeling and Difficulty Estimation for GUI Tasks

  • 将操作前的认知过程拆解为找、判、算等步骤,每步赋难度分
  • 通过大模型自动提取认知链,与用户完成时间相关性达0.46
  • 发现当前智能体在高认知难度任务上成功率显著下降

衡量图形界面任务难度对用户行为分析和智能体能力评估至关重要。现有基准多基于动作数量(如步骤数)量化难度,忽视了任务背后的心理负担。本文提出认知链框架,从认知视角建模任务难度:将执行前的认知过程分解为一系列认知步骤(如寻找、决策、计算),每步基于信息论赋予难度指数。我们开发基于大语言模型的方法,自动从任务执行轨迹中提取认知链。线性回归验证显示,所估认知难度与用户完成时间高度相关(标注后步骤级R²=0.46)。对前沿GUI智能体的评估表明,其在高认知难度任务上成功率显著降低,揭示出能力短板与人机一致性规律。最后讨论了该方法在智能体训练、能力评估及人机协同优化中的应用前景。

原文摘要 · Abstract (English)

Measuring GUI task difficulty is crucial for user behavior analysis and agent capability evaluation. Yet, existing benchmarks typically quantify difficulty based on motor actions (e.g., step counts), overlooking the cognitive demands underlying task completion. In this work, we propose Cognitive Chain, a novel framework that models task difficulty from a cognitive perspective. A cognitive chain decomposes the cognitive processes preceding a motor action into a sequence of cognitive steps (e.g., finding, deciding, computing), each with a difficulty index grounded in information theories. We develop an LLM-based method to automatically extract cognitive chains from task execution traces. Validation with linear regression shows that our estimated cognitive difficulty correlates well with user completion time (step-level R-square=0.46 after annotation). Assessment of state-of-the-art GUI agents shows reduced success on cognitively demanding tasks, revealing capability gaps and Human-AI consistency patterns. We conclude by discussing potential applications in agent training, capability assessment, and human-agent delegation optimization.

GUI任务认知建模难度评估智能体评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。