arXiv:2602.03059cs.HCcs.CL2026-02被引 1

语音指令变空间引导,让远程协助更直观高效

From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality

  • 仅凭语音识别目标,构建对象关系图实现空间定位
  • 实验显示任务效率提升,认知负荷降低,减少重复微调
  • 适合远程协作、无障碍指导等需要精准空间指引场景

我们提出Speech-to-Spatial,一种将口头远程协助指令转化为空间化增强现实(AR)引导的指代消解框架。不同于依赖手势、视线或人工标注的系统,该框架仅通过语音输入推断目标。基于对说话人指代模式的前期研究,我们归纳出四种常见指代方式:直接属性、关系型、记忆型与链式指代,并将其映射到以物体为中心的关系图中。给定一句话语,系统解析指代线索并生成持久的现场空间视觉提示,显著减少远程指导中反复微调(如“再往右一点”“现在停下”)的次数。我们在远程协助和意图消歧场景中验证了系统的有效性。评估结果表明,相比传统纯语音基线,Speech-to-Spatial提升了任务效率,降低了认知负担,增强了可用性,使无实体的语音指令转化为可在实时共享视图上可视化、可操作的引导。

原文摘要 · Abstract (English)

We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gesture, gaze) or manual expert annotations, Speech-to-Spatial infers the intended target solely from spoken references (speech input). Motivated by our formative study of speech referencing patterns, we characterize recurring ways people specify targets (Direct Attribute, Relational, Remembrance, and Chained) and ground them to our object-centric relational graph. Given an utterance, referent cues are parsed and rendered as persistent in-situ AR visual guidance, reducing iterative micro-guidance ("a bit more to the right", "now, stop.") during remote guidance. We demonstrate the use cases of our system with remote guided assistance and intent disambiguation scenarios. Our evaluation shows that Speechto-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to a conventional voice-only baseline, transforming disembodied verbal instruction into visually explainable, actionable guidance on a live shared view.

增强现实语音引导空间定位远程协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。