用大模型生成可理解的语义解释,让人工与多模态AI协作优化任务支持
ACE, Action and Control via Explanations: A Proposal for LLMs to Provide Human-Centered Explainability for Multimodal AI Assistants
- 通过大模型生成人类可懂的语义框架来解释行为
- 用户基于解释提供反馈,提升视觉、语音等模型的对齐精度
- 适合制造业人机协同场景,提升操作任务预测准确性
本文针对制造领域中多模态AI系统在人类绩效支持中的挑战展开研究。提出两个贡献:一是识别出参与式设计与训练中的关键难点;二是提出ACE范式——'通过解释实现行动与控制'。该范式利用大模型生成人类可理解的'语义框架'作为解释,使终端用户能据此提供所需数据,帮助对齐计算机视觉、自动语音识别及文档输入等多模态模型与表示。通过这种双向交互,人与AI共同构建对人类活动与行为的更准确建模,从而提升任务预测精度与支持效果,最终改善人工操作的任务表现。
原文摘要 · Abstract (English)
In this short paper we address issues related to building multimodal AI systems for human performance support in manufacturing domains. We make two contributions: we first identify challenges of participatory design and training of such systems, and secondly, to address such challenges, we propose the ACE paradigm: "Action and Control via Explanations". Specifically, we suggest that LLMs can be used to produce explanations in the form of human interpretable "semantic frames", which in turn enable end users to provide data the AI system needs to align its multimodal models and representations, including computer vision, automatic speech recognition, and document inputs. ACE, by using LLMs to "explain" using semantic frames, will help the human and the AI system to collaborate, together building a more accurate model of humans activities and behaviors, and ultimately more accurate predictive outputs for better task support, and better outcomes for human users performing manual tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。