人类协作中通过语言与手势形成抽象约定,提升效率。
Gesturing Toward Abstraction: Multimodal Convention Formation in Collaborative Physical Tasks
- 用语音和手势在物理任务中建立共享抽象规则。
- 多模态冗余使指令更准确,合作速度提升40%以上。
- 适合研究人机协作或具身智能系统的开发者。
人类智能的核心特征之一是通过重复协作动态形成临时约定以高效达成共同目标。本文通过在线单模态实验(n=98)研究自然语言如何构建抽象层级;后续实验室实验(n=40)则考察语音与手势在物理协作中的演化。参与者使用增强现实技术,一人观看3D虚拟塔并发出指令,另一人根据指令搭建实体塔。结果表明,双方通过建立语言与手势的抽象表达,并利用跨模态冗余强调关键变更,显著提升了执行速度与准确性。基于此,我们扩展了概率性约定形成模型至多模态场景,捕捉模态偏好变化。研究为设计具备约定意识的物理世界智能体提供了基础构件。
原文摘要 · Abstract (English)
A quintessential feature of human intelligence is the ability to create ad hoc conventions over time to achieve shared goals efficiently. We investigate how communication strategies evolve through repeated collaboration as people coordinate on shared procedural abstractions. To this end, we conducted an online unimodal study (n = 98) using natural language to probe abstraction hierarchies. In a follow-up lab study (n = 40), we examined how multimodal communication (speech and gestures) changed during physical collaboration. Pairs used augmented reality to isolate their partner's hand and voice; one participant viewed a 3D virtual tower and sent instructions to the other, who built the physical tower. Participants became faster and more accurate by establishing linguistic and gestural abstractions and using cross-modal redundancy to emphasize key changes from previous interactions. Based on these findings, we extend probabilistic models of convention formation to multimodal settings, capturing shifts in modality preferences. Our findings and model provide building blocks for designing convention-aware intelligent agents situated in the physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。