让对话导航理解‘我右边’变成‘东边’,提升复杂环境下的定位准确率。
Conversational Orientation Reasoning: Egocentric-to-Allocentric Navigation with Multimodal Chain-of-Thought
- 用多模态思维链分三步推理:提取空间关系、转绝对方向、推用户朝向。
- 语音识别后仍达98.1%准确率,比传统方法显著更优。
- 适合资源有限的场景,抗噪声和语言混杂能力强。
对话代理需将第一人称描述(如“在我右边”)转化为绝对方向(北/东/南/西)。这一挑战在室内或复杂设施中尤为关键,因GPS信号弱且缺乏详细地图。尽管思维链(CoT)提示已在语言与视觉任务中推动进展,但其在多模态空间方向推理中的应用仍不充分。本文提出对话方向推理(COR),一个面向中文真实场景的对话导航基准,涵盖非英语及语音识别转写情形。我们设计了多模态思维链(MCoT)框架,通过结构化三步流程整合语音识别转写文本与地标坐标:(1) 提取空间关系,(2) 将坐标映射为绝对方向,(3) 推断用户朝向。采用课程学习策略,在台湾-LLM-13B-v2.0-Chat这一中等规模模型上逐步构建能力,该模型代表资源受限设置。实验表明,MCoT在清晰转录下实现100%方向准确率,在含语音识别错误时仍达98.1%,显著优于单模态与非结构化基线。此外,MCoT在嘈杂对话条件下表现稳健,包括识别错误与多语言混用;跨域评估中也保持高精度,对语言变异、领域迁移和指代模糊具有鲁棒性。结果表明,结构化多模态思维链在可解释、资源高效具身导航中具备巨大潜力。
原文摘要 · Abstract (English)
Conversational agents must translate egocentric utterances (e.g., "on my right") into allocentric orientations (N/E/S/W). This challenge is particularly critical in indoor or complex facilities where GPS signals are weak and detailed maps are unavailable. While chain-of-thought (CoT) prompting has advanced reasoning in language and vision tasks, its application to multimodal spatial orientation remains underexplored. We introduce Conversational Orientation Reasoning (COR), a new benchmark designed for Traditional Chinese conversational navigation projected from real-world environments, addressing egocentric-to-allocentric reasoning in non-English and ASR-transcribed scenarios. We propose a multimodal chain-of-thought (MCoT) framework, which integrates ASR-transcribed speech with landmark coordinates through a structured three-step reasoning process: (1) extracting spatial relations, (2) mapping coordinates to absolute directions, and (3) inferring user orientation. A curriculum learning strategy progressively builds these capabilities on Taiwan-LLM-13B-v2.0-Chat, a mid-sized model representative of resource-constrained settings. Experiments show that MCoT achieves 100% orientation accuracy on clean transcripts and 98.1% with ASR transcripts, substantially outperforming unimodal and non-structured baselines. Moreover, MCoT demonstrates robustness under noisy conversational conditions, including ASR recognition errors and multilingual code-switching. The model also maintains high accuracy in cross-domain evaluation and resilience to linguistic variation, domain shift, and referential ambiguity. These findings highlight the potential of structured MCoT spatial reasoning as a path toward interpretable and resource-efficient embodied navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。