arXiv:2608.10651cs.RO2026-08

语音引导无人机时,人类会自然分成三个阶段,系统可据此自适应调整。

OAA: Three Phases of Vocal Guidance in Human-Drone Teleoperation

  • 通过轨迹变化点检测,自动识别语音引导的三阶段结构。
  • 三阶段中语音词汇类型与静默时间显著不同,统计检验差异极显著(p<0.001)。
  • 该结构在人-人与人-机场景中均出现,适合用于智能语音控制系统的自适应设计。

语音引导远程操作需要系统能适应人类引导行为的动态演变。然而,多数语音控制机器人系统将语音指令视为静态流,忽略了引导者在任务推进过程中沟通行为的变化。基于两组实验配置(人-人引导:10对,手指指指点;人-无人机遥控:29对,游戏手柄控制)的运动捕捉与语音数据,我们发现自发性语音引导始终可划分为三个在运动学与语言上截然不同的阶段:定向、接近和调整。这些阶段通过3D轨迹信号的变点检测自动识别,并经统计验证(Kruskal-Wallis检验,p<0.001)。三种词汇类别在两组配置中重复出现:旋转类词汇标记定向阶段,平移类词汇在此阶段稀少,而修饰语(attenuators)则在调整阶段累积。结合句间停顿,这些线索共同标定出定向阶段的边界,而仅靠语速无法实现。尽管操作接口差异巨大,三阶段结构在两种配置中均一致出现,表明其是人类空间引导的内在特征,而非实验设置的产物。本文讨论了该发现对语音引导远程操作中自适应控制的意义。

原文摘要 · Abstract (English)

Voice-guided teleoperation requires systems that adapt to the evolving dynamics of human guidance. Yet most voice-controlled robot systems treat spoken commands as a stationary stream, ignoring how the guide's communicative behavior changes as the task progresses. Using motion capture and speech data from two experimental configurations, humanhuman guidance (finger pointing, N =10 dyads) and humandrone teleoperation (gamepad control, N =29 dyads), we show that spontaneous vocal guidance consistently organizes into three kinematically and linguistically distinct phases: Orientation, Approach, and Adjustment. These phases are identified automatically via change point detection on 3D trajectory signals, and validated statistically (Kruskal-Wallis, p<.001). Three lexical families replicate across configurations: rotation vocabulary marks Orientation, translation vocabulary is scarce there, and attenuators accumulate toward Adjustment. Together with inter-utterance silence, these cues mark the Orientation boundary that speech rate alone leaves unmarked. The same three-phase structure emerges in both configurations despite radically different motor interfaces, suggesting it is an intrinsic property of human spatial guidance rather than an artifact of the experimental setup. We discuss implications for OAA-aware adaptive control in voice-guided teleoperation.

语音控制人机交互自适应系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。