arXiv:2605.01483cs.CVcs.AI2026-05

让工业机器人听懂复杂指令,提升人机协作可靠性

Research on Vision-Language Question Answering Models for Industrial Robots

  • 分层跨模态融合,整合视觉与语言信息进行联合推理
  • 在IVQA和RIF数据集上,准确率与鲁棒性显著优于现有模型
  • 适合需要精准理解操作指令的智能制造场景

针对现代制造中语义模糊、环境复杂及领域术语多等问题,提出一种分层跨模态融合模型用于工业机器人视觉-语言问答(VLQA)。该框架融合先进目标检测、多尺度视觉编码、句法解析与任务感知语义注意力,将视觉与语言信号统一到联合推理空间。基于区域的深度网络提取视觉特征,加权嵌入聚合,循环神经网络解析句子结构。通过自适应融合与交叉注意力驱动的细粒度语义对齐,系统可更可靠地处理操作查询、步骤指令与异常检测。在IVQA与RIF基准上的验证实验表明,该模型在语义对齐、Top-1准确率及对模糊或流程类查询的鲁棒性方面均有提升。消融实验证明多层级特征融合与上下文驱动门控对工业部署至关重要。本研究为提升工业机器人在多样化人机交互任务中的可解释性与执行效能提供了核心技术方法。

原文摘要 · Abstract (English)

A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of semantic ambiguity, complex environmental layouts, and domain-specific terminology common in modern manufacturing. The framework integrates advanced object detection, multi-scale visual encoding, syntactic parsing, and task-aware semantic attention to unite vision and language signals into a joint reasoning space. Region-based deep networks extract visual features, weighted embeddings aggregate, and recurrent neural parsing encodes sentence structures. Through fine-grained semantic alignment driven by adaptive fusion and cross-attention mechanisms, the system can handle operational queries, instruction steps, and anomaly detection with higher reliability. Compared to the existing VLQA benchmarks, validation experiments conducted on the IVQA and RIF benchmarks indicate improvements in semantic alignment, Top-1 accuracy, and robustness to ambiguous or procedural task queries. Ablation studies further quantify the impact of each architectural module, confirming the necessity of multi-level feature integration and context-driven gating for dependable industrial deployment. The technical advancements reported here provide core methodologies to improve the interpretability and operational effectiveness of industrial robots faced with diverse human-robot interaction tasks.

视觉语言问答工业机器人跨模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。