arXiv:2506.16931cs.AIcs.RO2025-06被引 1

用图像和图结构融合建模,实时求解机器人任务规划中的广义旅行商问题。

Multimodal Fused Learning for Solving the Generalized Traveling Salesman Problem in Robotic Task Planning

  • 将广义旅行商问题转化为坐标图像,结合图与视觉特征融合建模
  • 在多种实例上超越现有方法,且满足机器人实时计算需求
  • 适合需要高效路径规划的仓储、巡检等机器人应用场景

高效的任务规划对移动机器人至关重要,尤其在仓库拣选和环境监测等场景中。这类任务常需从多个目标簇中各选一个位置,形成广义旅行商问题(GTSP),其准确高效求解仍具挑战。为此,我们提出多模态融合学习(MMFL)框架,同时利用图与图像表示捕捉问题的互补信息,并学习生成高质量实时任务规划方案的策略。具体地,设计基于坐标的图像构建器,将GTSP实例转化为空间信息丰富的表征;引入自适应分辨率缩放策略以提升不同规模问题的适应性;构建带有专用瓶颈的多模态融合模块,实现几何与空间特征的有效整合。大量实验表明,所提方法在多种GTSP实例上显著优于当前最优方法,同时保持实时机器人应用所需的计算效率。物理机器人测试进一步验证了其在真实场景中的有效性。

原文摘要 · Abstract (English)

Effective and efficient task planning is essential for mobile robots, especially in applications like warehouse retrieval and environmental monitoring. These tasks often involve selecting one location from each of several target clusters, forming a Generalized Traveling Salesman Problem (GTSP) that remains challenging to solve both accurately and efficiently. To address this, we propose a Multimodal Fused Learning (MMFL) framework that leverages both graph and image-based representations to capture complementary aspects of the problem, and learns a policy capable of generating high-quality task planning schemes in real time. Specifically, we first introduce a coordinate-based image builder that transforms GTSP instances into spatially informative representations. We then design an adaptive resolution scaling strategy to enhance adaptability across different problem scales, and develop a multimodal fusion module with dedicated bottlenecks that enables effective integration of geometric and spatial features. Extensive experiments show that our MMFL approach significantly outperforms state-of-the-art methods across various GTSP instances while maintaining the computational efficiency required for real-time robotic applications. Physical robot tests further validate its practical effectiveness in real-world scenarios.

任务规划机器人图神经网络多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。