构建评估框架与数据集,验证AI辅助提升人类实操任务表现
Towards Effective Human-in-the-Loop Assistive AI Agents
- 设计多模态交互数据集与评估框架,量化人机协作效果
- 实测显示AI指导显著提升烹饪到战场急救等任务完成率
- 基于AR的智能助手适合医疗、制造等高风险实操场景
在日常活动和专业领域中,高效的人机协同完成物理任务具有巨大潜力。具备信息性指导能力的AI代理可提升人类表现,但因人机互动复杂性,评估仍具挑战。本文提出一种评估框架和多模态人机交互数据集,用于衡量AI指导对流程化任务表现、错误减少及学习成效的影响。同时开发了一款配备增强现实(AR)的AI代理,在烹饪到战场急救等真实任务中提供交互式指导。通过人机实验,揭示了AI辅助对人类表现的积极影响,并证明其能有效提升任务完成度。
原文摘要 · Abstract (English)
Effective human-AI collaboration for physical task completion has significant potential in both everyday activities and professional domains. AI agents equipped with informative guidance can enhance human performance, but evaluating such collaboration remains challenging due to the complexity of human-in-the-loop interactions. In this work, we introduce an evaluation framework and a multimodal dataset of human-AI interactions designed to assess how AI guidance affects procedural task performance, error reduction and learning outcomes. Besides, we develop an augmented reality (AR)-equipped AI agent that provides interactive guidance in real-world tasks, from cooking to battlefield medicine. Through human studies, we share empirical insights into AI-assisted human performance and demonstrate that AI-assisted collaboration improves task completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。