arXiv:2607.03163cs.RO2026-07

提出连续语义场,让机器人更稳定理解物体功能部位。

Beyond Point-Attached Semantics: Object-Centric Semantic Fields for Generalizable Manipulation

论文配图:Beyond Point-Attached Semantics: Object-Centric Semantic Fields for Generalizable Manipulation
图 1 · 摘自论文原文
  • 用连续语义场在3D任意位置读取语义特征,不依赖采样点。
  • 在仿真与真实双臂操作任务中,性能优于点云、2D投影等基线方法。
  • 适合需要泛化能力的机器人抓取与操作场景。

通用机器人操作需要对功能部件(如把手、工具头、开口和可抓区域)有稳定的三维理解。原始点云仅提供几何信息,缺乏显式的部件语义,且采样点随视角、传感器配置和物体实例变化。现有二维特征提升和离散三维点级特征虽增强了语义,但特征仍依附于观测依赖的采样点。本文提出一种以物体为中心的连续语义场,基于物体点云,在显式3D查询位置生成感知部件的语义嵌入。该场由标注部件的物体模型训练得到,随后冻结,作为操作策略的物体级条件生成语义点云。在RoboTwin仿真任务和真实世界双臂操作实验中,该表示提供更稳定的函数部件线索,显著提升策略性能,优于原始点云、2D特征提升和3D点级特征基线方法。

原文摘要 · Abstract (English)

Generalizable robot manipulation requires stable 3D understanding of functional object parts, such as handles, tool heads, openings, and graspable regions. Raw point clouds provide geometry but lack explicit part semantics, and their sampled points vary with viewpoint, sensor configuration, and object instance. Existing 2D feature lifting and discrete 3D point-wise features enrich point clouds with semantics, but the resulting features remain attached to observation-dependent samples. We propose an object-centric continuous semantic field that conditions on an object point cloud and reads part-aware semantic embeddings at explicit 3D query locations. The field is trained from part-annotated object models and then frozen to generate semantic point clouds as object-level conditioning for manipulation policies. Experiments on RoboTwin simulation tasks and real-world bimanual object manipulation show that our representation provides more stable functional-part cues and improves policy performance over raw point-cloud, 2D feature lifting, and 3D point-wise feature baselines. Project Page: \href{https://zainzh.github.io/beyond-point-attached-semantics}{https://zainzh.github.io/beyond-point-attached-semantics}.

机器人操作语义场点云理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。