arXiv:2605.11144cs.RO2026-05

让机器人提前预测动作完成后的3D状态,提升语言指令下抓取放置的成功率。

Forecast-aware Gaussian Splatting for Predictive 3D Representation in Language-Guided Pick-and-Place Manipulation

论文配图:Forecast-aware Gaussian Splatting for Predictive 3D Representation in Language-Guided Pick-and-Place Manipulation
图 1 · 摘自论文原文
  • 基于高斯点云生成任务完成时的3D预测状态,支持语言引导决策。
  • 真实场景测试中三项任务成功率分别达21/25、23/25、16/25。
  • 适合需要空间与语义推理的自主机器人操作研究者参考。

我们提出Forecast-aware Gaussian Splatting(Forecast-GS),一种面向语言引导抓取放置任务的预测性3D表示框架。现有系统多基于当前场景进行推理,未显式建模任务完成状态,这在部分观测下难以判断动作是否可达目标。Forecast-GS通过预测动作完成后环境的3D状态,提升动作可行性评估能力。我们在真实机器人上对Cutter-to-Box、Apple-to-Bowl、Sponge-to-Tray三项任务进行测试,每项任务在相同平台和感知设置下执行25次。使用自动候选选择时,成功率为21/25、23/25、16/25,优于ReKep基线(15/25、19/25、10/25)。人工辅助设定下成功率进一步提升至23/25、24/25、19/25,表明候选生成有效但自动排序仍有改进空间。结果说明显式预测终态可增强动作评估可靠性,而自动排序仍是实现完全自主操作的关键挑战。Forecast-GS为语言理解、3D感知与操作规划之间提供了可解释的桥梁。

原文摘要 · Abstract (English)

We introduce Forecast-aware Gaussian Splatting (Forecast-GS), a predictive 3D representation framework for language-conditioned robotic manipulation. While recent manipulation systems have made progress by grounding language instructions into robot affordances, value maps, or relational keypoint constraints, they usually reason over the current scene and do not explicitly model the task-completed state. This limitation is critical when success depends on satisfying spatial and semantic goals under partial observations, where the robot must evaluate whether a candidate action leads to a feasible task-consistent outcome. We validate Forecast-GS on real-world pick-and-place manipulation tasks, including Cutter-to-Box, Apple-to-Bowl, and Sponge-to-Tray. For each task, we conduct 25 real-world trials under varied initial object configurations using the same robot platform and sensing setup. Forecast-GS with automatic candidate selection achieves success rates of 21/25, 23/25, and 16/25 on the three tasks, respectively, outperforming the ReKep baseline, which achieves 15/25, 19/25, and 10/25. A diagnostic human-assisted setting further improves success rates to 23/25, 24/25, and 19/25, suggesting that candidate generation is effective while automatic ranking remains imperfect. These results suggest that explicitly forecasting task-completed 3D states enables more reliable action evaluation, while the gap between automatic and human-assisted selection indicates that robust final-state ranking remains an important challenge for fully autonomous manipulation. Overall, Forecast-GS provides an interpretable bridge between language understanding, 3D perception, and robotic manipulation planning.

3D表示机器人操作语言引导预测建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。