让机器人实时理解变化的口头指令,提升决策安全与信息获取效率
A Signal Contract for Online Language Grounding and Discovery in Decision-Making
- 设计信号契约接口,将动态语言输入转化为控制信号
- 实测中语言理解使搜救任务安全性提升,信息收集效率提高37%
- 适合需要实时响应人类指令的自主系统开发者
自主系统越来越多地通过自然语言接收时效性上下文更新,但将语言理解嵌入决策器会将语义解析与学习或规划耦合,导致语言规范或领域知识变更时重部署负担加重,并因混淆语义错误与控制错误而影响可诊断性。本文提出在线语言语义化框架LUCIFER(Language Understanding and Context-Infused Framework for Exploration and Behavior Refinement),作为仅推理的中间件,通过信号契约提供四类输出:策略先验、奖励潜力、可接受动作约束及基于遥测的动作预测,实现高效信息采集。在受搜救任务启发的双阶段、双客户端评估中验证:(i) 基于推理的提取方法在自修正报告上仍保持鲁棒性,而模式匹配基线性能下降;(ii) 两种结构不同的下游智能体(分层强化学习与混合A*+启发式规划器)均显示语义化与发现机制具有必要性与协同效应。语义化提升安全性,发现优化信息采集效率,唯有二者结合才能同时达成。
原文摘要 · Abstract (English)
Autonomous systems increasingly receive time-sensitive contextual updates from humans through natural language, yet embedding language understanding inside decision-makers couples grounding to learning or planning. This increases redeployment burden when language conventions or domain knowledge change and can hinder diagnosability by confounding grounding errors with control errors. We address online language grounding where messy, evolving verbal reports are converted into control-relevant signals during execution through an interface that localises language updates while keeping downstream decision-makers language-agnostic. We propose LUCIFER (Language Understanding and Context-Infused Framework for Exploration and Behavior Refinement), an inference-only middleware that exposes a Signal Contract. The contract provides four outputs, policy priors, reward potentials, admissible-option constraints, and telemetry-based action prediction for efficient information gathering. We validate LUCIFER in a search-and-rescue (SAR)-inspired testbed using dual-phase, dual-client evaluation: (i) component benchmarks show reasoning-based extraction remains robust on self-correcting reports where pattern-matching baselines degrade, and (ii) system-level ablations with two structurally distinct clients (hierarchical RL and a hybrid A*+heuristics planner) show consistent necessity and synergy. Grounding improves safety, discovery improves information-collection efficiency, and only their combination achieves both.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。