用自动对话生成扩充数据,让智能体导航更准更稳
Advancing DialNav through Automatic Embodied Dialog Augmentation

- 自动生成23.8万条多轮对话数据,解决训练样本少难题
- 在可见/不可见场景中成功率分别提升89%和100%
- 适合做具身智能、对话导航研究的学者参考
具备物理交互能力的智能体需具备生成与理解对话的能力,以确保安全与效率。尽管DialNav提供了一个在真实感室内环境中评估对话-执行闭环的框架,但其性能受限于训练数据稀缺(仅2000个任务)。为此,我们提出一种自动化生成流程,构建了大规模训练数据集RAINbow,包含23.8万个任务。该流程将现有视觉语言导航(VLN)数据转换为多轮对话,并实现高效高质量的数据生成。此外,我们引入两项互补改进:(1) 双策略训练,使导航训练与动态对话-导航循环对齐;(2) 利用VLN知识的定位模型。结合这些方法,我们的模型在 extbf{Val Seen}(58.24,+89\%)和 extbf{Val Unseen}(29.05,+100 extbackslash extbackslash%)两个测试集上的成功率显著超越基线,达到新基准。
原文摘要 · Abstract (English)
For embodied agents capable of physical interaction, the capability to create and understand dialog is crucial to ensure both safety and effectiveness. While DialNav~\cite{han2025dialnav} provides a framework for holistic evaluation of the dialog--execution loop in photorealistic indoor navigation, its performance remains limited by a critical scarcity of training data (2K episodes). To address this, we propose an automatic generation pipeline, and construct the \textbf{RAINbow} dataset, a large-scale training dataset with 238K episodes for DialNav. Our pipeline converts existing VLN datasets into multi-turn dialog and creates cost-efficient and high-quality dataset. Then, we introduce two additional complementary advances to unlock the data's full potential: (1) Dual-Strategy Training, a navigation training scheme to align the navigation training with the dynamic dialog-navigation loop, and (2) a localization model that leverages VLN knowledge. By combining these complementary solutions, our model substantially outperforms the baseline in success rate on both \textbf{Val Seen} (58.24, \textbf{+89\%}) and \textbf{Val Unseen} (29.05, \textbf{+100\%}) splits, establishing a new state of the art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。