实时追踪多人协作对话中的共同认知状态
TRACE: Real-Time Multimodal Common Ground Tracking in Situated Collaborative Dialogues
- 融合语音、动作、手势与视觉注意力,多模态追踪对话进展
- 动态更新群体对任务相关命题的认知立场与信念
- 适合需要实时理解协作意图的智能对话系统
我们提出TRACE,一种用于情境化协作对话中实时*共同认知*追踪的新系统。聚焦快速、实时性能,TRACE 跟踪参与者的话语、行为、手势及视觉注意力,利用这些多模态输入识别对话过程中提出的与任务相关的命题集合,并在任务推进过程中持续追踪群体对这些命题的元认知状态与信念变化。随着对能够促成协作的AI系统兴趣增加,TRACE为能参与多方、多模态对话的智能体提供了重要进展。
原文摘要 · Abstract (English)
We present TRACE, a novel system for live *common ground* tracking in situated collaborative tasks. With a focus on fast, real-time performance, TRACE tracks the speech, actions, gestures, and visual attention of participants, uses these multimodal inputs to determine the set of task-relevant propositions that have been raised as the dialogue progresses, and tracks the group's epistemic position and beliefs toward them as the task unfolds. Amid increased interest in AI systems that can mediate collaborations, TRACE represents an important step forward for agents that can engage with multiparty, multimodal discourse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。