给语言在机器人中的作用分类,并检验每种作用的证据是否扎实
What Language Does and What the Evidence Supports: A Functional Role Taxonomy and Evidence Audit of Language Grounding in Embodied Agents

- 提出语言在机器人中的五种功能角色,如指令指定、行为协调等
- 发现多数研究声称的语言作用缺乏可靠证据支持
- 用功能角色对比模型,避免过度推广结论
基础模型将语言广泛应用于具身智能体中,但其存在并不说明语言实际贡献了什么或贡献是否真正落地。本文通过分离这两个问题,定义了五种非互斥的语言功能角色:指令指定、具身表征、行为编排、接地调控和执行耦合。针对每种角色,追踪语言内容到具身实体的路径,并识别可验证责任的观测或干预手段。将该框架应用于文献分析后发现,功能使用与证据支持之间存在普遍差距:可解释或修改后的语言中间表示可能错误、未被使用,或无法影响后续行为。即使动作直接由语言决定,系统整体成功也无法独立证明语言的作用。因此,我们逐条评估接地主张,考察报告证据是否支持所赋予的语言责任。以功能角色而非架构为比较单位,使模块化与端到端智能体可比,且不超出实证范围。
原文摘要 · Abstract (English)
Foundation models place language throughout embodied agents, but its presence does not show what it contributes or how well that contribution is grounded. This survey separates these two questions. We define five non-exclusive functional roles for language: Specification, Embodied Representation, Action Orchestration, Grounding Regulation, and Execution Coupling. For each role, we trace the path from linguistic content to its embodied consumer and identify the observations or interventions that can test the claimed responsibility. Applying this framework to the reviewed literature reveals a recurring gap between functional use and evidential support. Interpretable or revised linguistic intermediates may be incorrect, go unused, or fail to affect later behavior. Even when actions are directly conditioned on language, system-level success does not by itself isolate language's contribution. We therefore evaluate grounding claim by claim, asking whether the reported evidence supports the specific responsibility assigned to language. Using role claims rather than architectures as the unit of comparison allows us to compare modular and end-to-end embodied agents without extending conclusions beyond the reported evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。