arXiv:2607.13621cs.AI2026-07

构建首个统一的寻人跟随基准,解决先找后跟的现实挑战

UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

论文配图:UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
图 1 · 摘自论文原文
  • 提出任务驱动路由机制,动态切换寻人与跟随模式
  • 在单人/多人环境中显著优于单头与双头基线模型
  • 适合研究具身智能、多模态导航与语言引导行为的学者

语言引导的人类跟随是具身智能体的重要能力,但现有基准通常假设目标人物在任务开始时可见,忽略了更现实的场景:智能体需先根据语言描述寻找目标,再在动态环境中持续跟随。现有研究多限于特定任务场景,依赖较强环境先验知识,且将搜索与跟随视为独立任务,缺乏统一评估框架。为此,我们提出统一具身寻人与跟随基准(UESF-Bench),涵盖大规模多样场景,要求智能体具备语义引导探索、行为切换与恢复、延迟身份识别等能力。为此设计了SeekFollow-VLA框架,采用视觉-语言-动作一体化结构,结合任务驱动路由机制实现隐状态推理与状态转换建模。实验表明,该方法在单人与多人环境中均显著优于单头与双头基线,建立了统一寻人跟随任务的新基准。

原文摘要 · Abstract (English)

Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visible at the start of an episode. This setting simplifies the problem and overlooks a more realistic requirement: an agent often needs to first find a language-described target and then persistently follow that target in a dynamic environment. While recent work has started to study human search, existing settings are typically evaluated in task-specific scenarios and often rely on stronger prior knowledge of the environment. Moreover, they usually treat searching and following as separate tasks and still lack a unified benchmark for systematic evaluation. To address these limitations, we introduce the Unified Embodied Seeking and Following Benchmark (UESF-Bench), a large-scale and diverse benchmark for embodied human seeking and following. The benchmark requires agents to handle semantic-guided exploration, reliable behavior switching and recovery, and delayed identity grounding. To this end, we propose SeekFollow-VLA, a vision-language-action framework with a task-driven routing mechanism for latent phase inference and transition modeling between seeking and following. Experimental results show that SeekFollow-VLA achieves clear improvements over both single-head and dual-head baselines across single-person and multi-person environments, establishing a baseline for unified embodied seek-and-follow.

具身智能语言导航多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。