从视觉图像自动发现机器人重排任务的抽象结构。
Learning Discrete Abstractions for Visual Rearrangement Tasks Using Vision-Guided Graph Coloring
- 用视觉距离与结构约束结合,生成图结构抽象
- 在模拟任务中识别出有效抽象,提升规划效率
- 适合需要自主理解场景的智能机器人研究者
从数据中学习抽象是机器人领域的核心挑战。人类天然以高层目标进行推理,将执行交由底层动作技能——这种能力使复杂环境中的高效求解成为可能。在机器人领域,抽象与分层推理长期用于规划,但通常依赖人工设计,耗时且难扩展。若能直接从视觉数据中自动发现有用抽象,将显著提升规划系统的可扩展性与实际应用能力。本文聚焦于状态由原始图像表示的重排任务,提出一种方法:通过结合结构约束与注意力引导的视觉距离,生成离散的图结构抽象。该方法利用重排问题的固有二分结构,将结构约束与视觉嵌入统一建模,实现仅凭视觉信息自主发现抽象,进而支持高层规划。我们在两个模拟重排任务上评估该方法,结果表明其能持续识别出有意义的抽象,显著提升规划效果,并优于现有方法。
原文摘要 · Abstract (English)
Learning abstractions directly from data is a core challenge in robotics. Humans naturally operate at an abstract level, reasoning over high-level subgoals while delegating execution to low-level motor skills -- an ability that enables efficient problem solving in complex environments. In robotics, abstractions and hierarchical reasoning have long been central to planning, yet they are typically hand-engineered, demanding significant human effort and limiting scalability. Automating the discovery of useful abstractions directly from visual data would make planning frameworks more scalable and more applicable to real-world robotic domains. In this work, we focus on rearrangement tasks where the state is represented with raw images, and propose a method to induce discrete, graph-structured abstractions by combining structural constraints with an attention-guided visual distance. Our approach leverages the inherent bipartite structure of rearrangement problems, integrating structural constraints and visual embeddings into a unified framework. This enables the autonomous discovery of abstractions from vision alone, which can subsequently support high-level planning. We evaluate our method on two rearrangement tasks in simulation and show that it consistently identifies meaningful abstractions that facilitate effective planning and outperform existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。