arXiv:2608.25623cs.AIcs.CY2026-08

用认知能力画像评估AI在工作中的适用性,帮企业决定哪些任务该自动化。

Using profiles of cognitive capability to assess AI suitability for workplace tasks

论文配图:Using profiles of cognitive capability to assess AI suitability for workplace tasks
图 1 · 摘自论文原文
  • 通过认知能力维度统一评估AI与任务需求
  • 6个AI系统在不同认知维度表现差异大,非模型家族决定
  • 适合组织规划AI试点,也适用于人机协作设计

组织在部署AI时面临难题:哪些任务可自动化、哪些应由人完成、哪些适合人机协作。通用基准分数难以预测实际表现,而对模型能力的人工判断又迅速过时。本文提出一种管道,通过一组共享的核心认知能力对智能体和任务进行画像。认知能力画像基于带认知需求标注的基准测试集,推断智能体的能力;任务需求加权则从领域专家处获取各项能力在具体工作中的相对重要性。两者使用相同认知维度,可独立更新,并组合估算在领域、组织、岗位或具体职责层面的AI适用性。我们验证了合成智能体的能力恢复效果,对6个AI系统进行了画像,并从6个职业领域的410名员工中收集任务需求。结果表明,不同AI系统在认知维度上的差异大于模型家族间的差异,而各类工作活动趋同于一个共享的认知核心。生成的评分可作为比较工具,识别适合试点的任务及当前系统不匹配的领域。我们进一步讨论将框架扩展至画像人类工作者,实现从AI适用性向人机任务分配的演进。

原文摘要 · Abstract (English)

Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where systems will succeed or fail in practice, while human judgements of model capabilities quickly become outdated. We introduce a pipeline that profiles agents and tasks using a shared set of core cognitive capabilities. Cognitive capability profiling infers an agent's capabilities from performance on a benchmark battery annotated for the cognitive demands of each item. Task requirements weighting elicits from domain experts the relative importance of these same capabilities for their work. As both use a common set of cognitive dimensions, they can be updated independently as models and roles change, and combined to estimate AI suitability at the level of a domain, organisation, role, or individual duty. We validate capability recovery on synthetic agents, profile six AI systems, and elicit task requirements from 410 employees across six occupational domains. AI systems differed more across cognitive dimensions than across model families, while workplace activities converged on a shared cognitive core. The resulting scores provide a comparative scoping tool for identifying promising candidates for piloting and areas where current systems are unlikely to be well suited. We discuss extending the framework to profile human workers alongside AI systems, moving from AI suitability towards human-machine task allocation.

AI评估人机协作认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。