arXiv:2512.17992cs.RO2025-12被引 4

用大模型统一生成与学习机器人谓词,提升复杂任务成功率

Unifying Deep Predicate Invention with Pre-trained Foundation Models

  • 双层学习框架:大模型生成谓词假设,神经网络从数据中学习并反馈优化
  • 在5个仿真和1个真实机器人场景中,成功率比纯自上而下方法高2-4倍
  • 适合需要灵活符号建模的智能机器人系统,尤其擅长杂乱环境下的推理

长时序机器人任务因连续状态动作空间和稀疏反馈而困难。符号世界模型通过将任务分解为描述物体属性与关系的离散谓词来缓解此问题。现有方法要么自上而下地仅靠提示预训练模型而无数据支撑,要么自下而上地从示范中学习但缺乏高层先验。本文提出UniPred,一种统一两种路径的双层学习框架。UniPred利用大语言模型(LLMs)生成谓词效果分布,监督神经谓词学习;学习到的反馈则迭代优化LLM假设。结合强大的视觉基础模型特征,UniPred在杂乱场景中学习出鲁棒的谓词分类器。我们还提出一种谓词评估方法,支持超越STRIPS假设的符号模型。在五个仿真和一个真实机器人领域中,UniPred的成功率比自上而下方法高2-4倍,学习速度比自下而上方法快3-4倍,推动了可扩展、灵活的机器人符号世界建模。

原文摘要 · Abstract (English)

Long-horizon robotic tasks are hard due to continuous state-action spaces and sparse feedback. Symbolic world models help by decomposing tasks into discrete predicates that capture object properties and relations. Existing methods learn predicates either top-down, by prompting foundation models without data grounding, or bottom-up, from demonstrations without high-level priors. We introduce UniPred, a bilevel learning framework that unifies both. UniPred uses large language models (LLMs) to propose predicate effect distributions that supervise neural predicate learning from low-level data, while learned feedback iteratively refines the LLM hypotheses. Leveraging strong visual foundation model features, UniPred learns robust predicate classifiers in cluttered scenes. We further propose a predicate evaluation method that supports symbolic models beyond STRIPS assumptions. Across five simulated and one real-robot domains, UniPred achieves 2-4 times higher success rates than top-down methods and 3-4 times faster learning than bottom-up approaches, advancing scalable and flexible symbolic world modeling for robotics.

机器人符号学习大模型谓词发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。