arXiv:2603.05230cs.CVcs.RO2026-03

用数字孪生+视觉语言模型,实现自动分拣衣物与异物检测。

Digital Twin Driven Textile Classification and Foreign Object Recognition in Automated Sorting Systems

  • 构建数字孪生系统,融合多模态感知与语义推理。
  • Qwen模型准确率达87.9%,在异物识别上表现优异。
  • 适合工业级衣物分拣场景,支持边缘部署轻量模型。

可持续纺织品回收需求推动自动化分拣技术发展,尤其需应对可变形衣物与杂乱环境中的异物检测。本文提出一种基于数字孪生的双臂机器人分拣系统,集成抓取预测、多模态感知与语义推理,实现真实世界衣物分类。该系统配备RGBD传感器、电容式触觉反馈及碰撞感知运动规划,能自主从无序篮中分拣衣物,转移至检测区,并利用前沿视觉语言模型(VLMs)进行分类。在包含223个检测场景的数据集上评估了九种来自五个模型家族的VLM,涵盖衬衫、袜子、裤子、内衣、异物(非上述类别的衣物)及空场景。评估指标包括各类别准确率、幻觉行为及实际硬件约束下的计算性能。结果表明,Qwen模型家族整体准确率最高(达87.9%),异物检测能力突出;轻量级模型如Gemma3则在边缘部署中展现良好速度-精度平衡。数字孪生结合MoveIt实现碰撞感知路径规划,并将检测后衣物的分割3D点云整合至虚拟环境,提升操作可靠性。本系统验证了将语义VLM推理与传统抓取检测、数字孪生技术融合,在真实工业场景中实现可扩展、自主纺织品分拣的可行性。

原文摘要 · Abstract (English)

The increasing demand for sustainable textile recycling requires robust automation solutions capable of handling deformable garments and detecting foreign objects in cluttered environments. This work presents a digital twin driven robotic sorting system that integrates grasp prediction, multi modal perception, and semantic reasoning for real world textile classification. A dual arm robotic cell equipped with RGBD sensing, capacitive tactile feedback, and collision-aware motion planning autonomously separates garments from an unsorted basket, transfers them to an inspection zone, and classifies them using state of the art Visual Language Models (VLMs). We benchmark nine VLM s from five model families on a dataset of 223 inspection scenarios comprising shirts, socks, trousers, underwear, foreign objects (including garments outside of the aforementioned classes), and empty scenes. The evaluation assesses per class accuracy, hallucination behavior, and computational performance under practical hardware constraints. Results show that the Qwen model family achieves the highest overall accuracy (up to 87.9 %), with strong foreign object detection performance, while lighter models such as Gemma3 offer competitive speed accuracy trade offs for edge deployment. A digital twin combined with MoveIt enables collision aware path planning and integrates segmented 3D point clouds of inspected garments into the virtual environment for improved manipulation reliability. The presented system demonstrates the feasibility of combining semantic VLM reasoning with conventional grasp detection and digital twin technology for scalable, autonomous textile sorting in realistic industrial settings.

数字孪生衣物分拣视觉语言模型异物检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。