arXiv:2507.04351cs.ROcs.AI2025-07中稿 · ICRA被引 1

用多模态大模型实现布料智能分类与选择,提升纺织机器人精度。

MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection

  • 基于多模态大模型与视觉触觉数据,实现布料属性排序。
  • 在220种布料上验证,性能优于现有视觉语言模型。
  • 适合智能制造、智能零售领域研究者参考。

在机器人纺织制造、服装生产和智能零售中,选择合适的布料对满足功能与质量要求至关重要。我们提出MLLM-Fabric,一个基于多模态大语言模型(MLLM)的布料分类与选择机器人框架。该系统构建于多模态机器人平台之上,通过监督微调和解释引导的蒸馏方法训练,用于布料属性排序。我们还发布了包含220种多样布料的数据集,每种布料均配有RGB图像以及同步的视觉-触觉和压力数据。实验表明,Fabric-Llama-90B在属性排序和选择可靠性方面均持续优于预训练的视觉-语言基线模型。代码与数据集已公开:https://github.com/limanwang/MLLM-Fabric。

原文摘要 · Abstract (English)

Choosing appropriate fabrics is critical for meeting functional and quality demands in robotic textile manufacturing, apparel production, and smart retail. We propose MLLM-Fabric, a robotic framework leveraging multimodal large language models (MLLMs) for fabric sorting and selection. Built on a multimodal robotic platform, the system is trained through supervised fine-tuning and explanation-guided distillation to rank fabric properties. We also release a dataset of 220 diverse fabrics, each with RGB images and synchronized visuotactile and pressure data. Experiments show that our Fabric-Llama-90B consistently outperforms pretrained vision-language baselines in both attribute ranking and selection reliability. Code and dataset are publicly available at https://github.com/limanwang/MLLM-Fabric.

多模态机器人布料识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。