用神经符号框架解决机器人装配中的语义与物理执行脱节问题
Bridging Semantics and Physical Execution: A Neuro-Symbolic Framework for Multi-Pair Robotic Assembly

- 分层构建每对组件的最优子图,用大模型生成基础动作避免幻觉
- 在100个真实场景中实现97%全局可执行性,实机部署成功率90%
- 适合需要高可靠性的复杂自主装配任务,尤其抗干扰场景
非结构化环境中多对机器人装配面临空间干扰和接触不确定性。现有方法难以衔接认知决策与物理执行,或因状态空间爆炸与知识瓶颈,或出现逻辑幻觉与拓扑冲突。本文提出端到端神经符号框架:先为每对组件生成最优子图,解耦通用性与边缘情况,再化解跨对干扰。基于眼在手上RGB-D场景,框架提取语义实例身份与状态,并量化场景以计算发散度。每对组件通过大语言模型生成基础动作,减少幻觉;轻量级判别器推理并插入支持性动作应对边缘情况。基于量化基线与当前场景的发散度,系统可低成本扩展。增强后的子图经拓扑协调形成全局序列,保持行为一致性。嵌入原子技能的动态行为树闭环控制力觉执行。离线评估在100个真实场景中达到97.00%全局可执行性,优于经典与最先进规划器。实机部署于UR3机械臂,在强干扰下实现90%成功率,定位精度达0.5 mm,验证了复杂自主装配的统一且可验证解决方案。
原文摘要 · Abstract (English)
Multi-pair robotic assembly in unstructured environments faces spatial interference and contact uncertainties. Existing paradigms fail to bridge cognitive decision-making and physical execution, as they either encounter state-space explosion and knowledge bottlenecks or suffer from logical hallucinations and topological conflicts. We propose an end-to-end neuro-symbolic framework that solves the challenge hierarchically: generating optimal subgraphs for each pair, decoupling generality from edge cases, and then resolving cross-pair interferences. Given an eye-on-hand RGB-D assembly scene, the framework extracts semantic instance identity and state while quantifying the scene for divergence calculation. For each pair, optimal subgraph is generated via LLM using barely basic actions to mitigate hallucinations. Supportive actions for edge cases are reasoned and inserted with a lightweight discriminator. Driven by the divergence between the quantified baseline and current scene, it is easily extensible at low cost. Augmented subgraphs are topologically coordinated into global sequences while preserving internal behavioral coherence. Dynamic behavior trees embedding atomic skills close the force-aware execution loop. Offline evaluation on 100 real-world scenes achieves 97.00% global executability, outperforming classical and state-of-the-art planners. Real-robot deployment on a UR3 arm attains 90% success rate with 0.5 mm tolerance under strong interference, demonstrating a unified and verifiable solution for complex autonomous assembly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。