arXiv:2608.11692cs.AI2026-08

提升视觉语言模型在多视角物流分拣中的规划能力

HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting

论文配图:HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
图 1 · 摘自论文原文
  • 通过内生数据增强和全局上下文排序,解决跨视角监督不足问题
  • 在SortingBench上使Qwen3-VL-8B准确率从63.6%提升至78.8%
  • 适用于工业级自主分拣系统,支持真实场景部署

自主物流分拣系统(ALSS)是具身AI的重要工业应用,需对空间分离的摄像头视图进行联合规划。本文将该场景建模为联合多场景理解(JMSU)。具备开放世界视觉理解和任务规划能力的视觉语言模型(VLM)是解决JMSU的有力候选,但直接应用现有VLM存在跨场景监督稀疏和长视觉上下文导致注意力分散的问题。为此,本文提出HUGIN训练框架,包含两个互补组件:内生数据增强在操作约束下重组已验证的原子事实;全局上下文排序强化指令表征与完整视觉上下文而非部分上下文的对齐。为支持研究,构建了来自四个布局的高质量工业分拣数据集与基准SortingBench。在五个开源VLM上,HUGIN一致优于基线;例如,Qwen3-VL-8B在SortingBench上的准确率从63.6%提升至78.8%。额外实验验证了各组件有效性及JMSU在具身任务中的溢出效益。超过15,000个包裹的部署测试证实了基于VLM规划在实际场景中的可行性。

原文摘要 · Abstract (English)

Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spatially disjoint camera views. We formulate this setting as Joint Multi-Scene Understanding (JMSU). With open-world visual understanding and task-planning capabilities, vision-language models (VLMs) are promising candidates for JMSU. However, directly applying existing VLMs to JMSU is non-trivial due to scarce cross-scene supervision and attention dispersion caused by long visual context in JMSU. To address these challenges, we propose HUGIN, a training framework with two complementary components. Endogenous Data Augmentation recombines verified atomic facts under operating constraints, while Global Context Ranking aligns the instruction representation more strongly with the complete visual context than with a partial visual context. To support ongoing research, we construct a high-quality industrial sorting dataset and benchmark named SortingBench from four layouts of autonomous logistics sorting systems. Across five open VLMs, HUGIN consistently outperforms matched baselines; for example, the accuracy on SortingBench of Qwen3-VL-8B increases from 63.6% to 78.8%. Additional experiments verify the effectiveness of each component and JMSU's spillover benefits in embodied tasks. Deployment tests involving more than 15,000 packages support the practical viability of VLM-based planning for autonomous logistics sorting.

视觉语言物流分拣具身智能多视角理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。