arXiv:2505.11962cs.AI2025-05ACL被引 2

构建复杂多模态指令跟随基准,评估模型在动态环境中的理解与适应能力

CrafText Benchmark: Advancing Instruction Following in Complex Multimodal Open-Ended World

  • 设计包含3924条指令的多模态测试集,覆盖定位、条件、建造和成就任务
  • 使用3423个独特词汇,检验模型对复杂语言和新指令形式的泛化能力
  • 适合研究多模态推理、指令理解与自适应决策的开发者和研究人员

在真实世界中遵循指令需要应对环境的动态性与复杂性:环境不断变化且不可预测,指令语言多样、词汇丰富,可能的目标数量庞大。尽管该领域已有大量研究,但多数实验仍基于静态环境、简单指令和有限词汇,难以评估智能体在更复杂场景下的表现。为填补这一空白,我们提出CrafText,一个用于评估多模态环境中指令跟随能力的基准。CrafText包含3,924条指令,涉及3,423个唯一词汇,涵盖定位、条件、建造和成就四类任务。我们还设计了一套评估协议,衡量智能体对新型指令表达和动态任务配置的泛化能力,从而严格测试其语言理解与自适应决策水平。

原文摘要 · Abstract (English)

Following instructions in real-world conditions requires the ability to adapt to the world's volatility and entanglement: the environment is dynamic and unpredictable, instructions can be linguistically complex with diverse vocabulary, and the number of possible goals an agent may encounter is vast. Despite extensive research in this area, most studies are conducted in static environments with simple instructions and a limited vocabulary, making it difficult to assess agent performance in more diverse and challenging settings. To address this gap, we introduce CrafText, a benchmark for evaluating instruction following in a multimodal environment with diverse instructions and dynamic interactions. CrafText includes 3,924 instructions with 3,423 unique words, covering Localization, Conditional, Building, and Achievement tasks. Additionally, we propose an evaluation protocol that measures an agent's ability to generalize to novel instruction formulations and dynamically evolving task configurations, providing a rigorous test of both linguistic understanding and adaptive decision-making.

指令跟随多模态基准测试泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。