构建可交互的闭环模拟器,让智能体学会像人一样提问和理解导航指引。
FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI

- 用大模型实现高阶行为规划与语义化交互,结合意图和社会认知理论。
- 生成63,429帧标注数据、超17小时交互记录,支持复杂任务训练与评估。
- 适合研究人机互动、具身智能的开发者,尤其关注自然对话与导航的场景。
随着具身智能成为人工智能核心前沿,仿真平台需超越底层物理交互,捕捉复杂的人类中心社交行为。我们提出FreeAskWorld,一个集成大语言模型(LLMs)的交互式仿真框架,用于高阶行为规划与语义化交互,基于意图与社会认知理论。该框架支持可扩展、逼真的人-智能体仿真,并配备模块化数据生成管道,适配多样具身任务。为验证框架,我们将经典视觉-语言导航(VLN)任务拓展为包含主动询问的交互式方向问询场景,使智能体能主动寻求并解读导航指导。我们公开发布FreeAskWorld,一个大规模基准数据集,包含重构环境、六种任务类型、16类核心物体、63,429个标注样本帧及超过17小时交互数据,以支持具身AI系统的训练与评估。我们在开环与闭环设置下对VLN模型与人类参与者进行基准测试,结果表明,经FreeAskWorld微调的模型在语义理解与交互能力上显著优于原模型。这证明了基于社会认知的仿真框架对推动具身智能向高级规划与自然人机交互发展的有效性。重要的是,交互本身已成为一种额外的信息模态。
原文摘要 · Abstract (English)
As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex, human-centered social behaviors. We introduce FreeAskWorld, an interactive simulation framework that integrates large language models (LLMs) for high-level behavior planning and semantically grounded interaction, informed by theories of intention and social cognition. Our framework supports scalable, realistic human-agent simulations and includes a modular data generation pipeline tailored for diverse embodied tasks. To validate the framework, we extend the classic Vision-and-Language Navigation (VLN) task into a interaction enriched Direction Inquiry setting, wherein agents can actively seek and interpret navigational guidance. We present and publicly release FreeAskWorld, a large-scale benchmark dataset comprising reconstructed environments, six diverse task types, 16 core object categories, 63,429 annotated sample frames, and more than 17 hours of interaction data to support training and evaluation of embodied AI systems. We benchmark VLN models, and human participants under both open-loop and closed-loop settings. Experimental results demonstrate that models fine-tuned on FreeAskWorld outperform their original counterparts, achieving enhanced semantic understanding and interaction competency. These findings underscore the efficacy of socially grounded simulation frameworks in advancing embodied AI systems toward sophisticated high-level planning and more naturalistic human-agent interaction. Importantly, our work underscores that interaction itself serves as an additional information modality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。