arXiv:2604.17019cs.AI2026-04

研究指令粒度对智能体行为的影响,发现精细与粗略指令效果最好。

Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents

论文配图:Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents
图 1 · 摘自论文原文
  • 构建多粒度指令基准Mini-BEHAVIOR-Gran,支持同一任务不同描述层次测试。
  • 发现指令粒度与性能呈U型关系,极细和极粗指令下表现最优。
  • 粗粒度优势源于视觉主导策略,适合研究视觉感知强化的场景。

指令粒度是语言引导具身智能领域中一个重要但控制不足的变量。现有基准通常为每个任务配置单一固定指令,难以研究相同任务在不同描述细节下的行为变化。本文提出Mini-BEHAVIOR-Gran,作为Mini-BEHAVIOR的扩展,为每个任务提供从高层目标到分步指导的多种指令变体。通过该基准,我们对比了四种跨任务粒度量化指标:词元数量、实体数、动作动词数与规划宽度,发现规划宽度与智能体性能相关性最强。进一步以宽度组织训练与评估,揭示出指令粒度与性能之间非单调的U型关系,在精细与粗略两个极端均出现性能峰值。深入分析表明,粗粒度下的性能回升与浅层语义接地有关,即智能体学习到以视觉为主的策略。

原文摘要 · Abstract (English)

Instruction granularity is an important yet poorly controlled variable in language-guided embodied AI. Existing benchmarks typically pair each task with a single static instruction, making it difficult to study how agent behavior changes when the same task is described at different levels of detail. We introduce Mini-BEHAVIOR-Gran, a new benchmark for controlled studies of instruction granularity that extends Mini-BEHAVIOR with multiple instruction variants per task, ranging from high-level goal descriptions to step-by-step guidance. Using this benchmark, we compare four candidate metrics for cross-task granularity quantification: token count, entity count, action-verb count, and planning-width, and find that width correlates most consistently with agent performance. Using width to organize training and evaluation further reveals a non-monotonic U-shaped relationship between instruction granularity and performance, with peaks at both fine and coarse extremes. Further analysis suggests that the coarse-granularity performance rebound is associated with shallow grounding, where agents learn vision-dominant policies.

具身智能指令粒度行为建模视觉策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。