arXiv:2509.23121cs.AI2025-09中稿 · IAI 2025被引 5

评估视觉语言动作模型在工业场景中的表现与挑战

Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges

  • 对比主流VLA模型在工业场景下的性能表现
  • 细调后可完成简单抓取,复杂任务仍不理想
  • 适合关注AI工业落地的工程师与研究者

人工智能在工业领域的应用正推动传统自动化向具备感知与认知能力的智能系统转型。视觉语言动作(VLA)模型作为统一感知、推理与控制的关键范式,其性能是否满足工业需求?本文从工业部署角度,对比现有最先进VLA模型在工业场景中的表现,并从数据采集与模型架构视角分析其实际部署的局限性。结果表明,经微调后VLA模型在工业环境中仍能完成简单抓取任务;但在复杂环境、多样物体类别及高精度放置任务中仍有显著提升空间。研究为VLA模型在工业应用中的适应性提供了实践洞察,强调需通过任务定制化改进以增强其鲁棒性、泛化性与精度。

原文摘要 · Abstract (English)

The application of artificial intelligence (AI) in industry is accelerating the shift from traditional automation to intelligent systems with perception and cognition. Vision language-action (VLA) models have been a key paradigm in AI to unify perception, reasoning, and control. Has the performance of the VLA models met the industrial requirements? In this paper, from the perspective of industrial deployment, we compare the performance of existing state-of-the-art VLA models in industrial scenarios and analyze the limitations of VLA models for real-world industrial deployment from the perspectives of data collection and model architecture. The results show that the VLA models retain their ability to perform simple grasping tasks even in industrial settings after fine-tuning. However, there is much room for performance improvement in complex industrial environments, diverse object categories, and high precision placing tasks. Our findings provide practical insight into the adaptability of VLA models for industrial use and highlight the need for task-specific enhancements to improve their robustness, generalization, and precision.

视觉语言动作工业AI模型部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。