用统一建模语言让机器人更懂清理房间的逻辑和步骤
UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
- 用UML类图和活动图构建可执行的推理与规划
- 在MRoom-30k数据集上执行成功率显著提升
- 适合需要高可解释性与可执行性的机器人任务
思维链(CoT)提示提升了大语言模型的推理能力,但其依赖非结构化文本,在具身任务中限制了可解释性和可执行性。以往研究尝试使用场景图或逻辑图实现结构化思维链,但仍存在仅能表达低阶关系、缺乏继承或行为抽象等构造、且无标准化语义支持顺序或条件规划等问题。本文提出UML-CoT,利用统一建模语言(UML)生成符号化思维链与可执行动作计划:类图捕获对象的组合语义,活动图建模过程控制流。通过三阶段训练流程,结合监督微调与组相对策略优化(GRPO),并利用仅答案数据进行奖励学习。在新提出的杂乱房间清洁场景基准MRoom-30k上评估,UML-CoT在可解释性、规划连贯性和执行成功率方面均优于非结构化思维链,证明了UML作为更丰富、更具行动力的结构化推理形式主义的有效性。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) prompting improves reasoning in large language models (LLMs), but its reliance on unstructured text limits interpretability and executability in embodied tasks. Prior work has explored structured CoTs using scene or logic graphs, yet these remain fundamentally limited: they model only low-order relations, lack constructs like inheritance or behavioral abstraction, and provide no standardized semantics for sequential or conditional planning. We propose UML-CoT, a structured reasoning and planning framework that leverages Unified Modeling Language (UML) to generate symbolic CoTs and executable action plans. UML class diagrams capture compositional object semantics, while activity diagrams model procedural control flow. Our three-stage training pipeline combines supervised fine-tuning with Group Relative Policy Optimization (GRPO), including reward learning from answer-only data. We evaluate UML-CoT on MRoom-30k, a new benchmark of cluttered room-cleaning scenarios. UML-CoT outperforms unstructured CoTs in interpretability, planning coherence, and execution success, highlighting UML as a more expressive and actionable structured reasoning formalism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。