提出新剪枝方法,让机器人视觉模型既快又可信。
Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding

- 基于推理依据优化剪枝,让证据与判断更匹配。
- 在真实视频数据上准确率最高,双正确预测率领先。
- 适合需要安全可靠决策的机器人交互场景。
眼动视觉理解中视觉语言模型(VLMs)的快速崛起,使得人-机器人协作任务中的低延迟推理愈发关键。为满足机载处理和实时交互机器人的效率需求,可直接应用针对VLMs的权重剪枝技术以缩小模型规模和计算量。此外,安全的人机交互要求剪枝策略能保留双重正确预测:输出既需准确,又要有确凿证据支撑,以降低风险并保障用户信任。本文从双重正确预测视角重新审视VLM剪枝。实验令人意外地发现,现有剪枝方法常保留正确的证据定位,却削弱了正确预测能力。为此,我们提出一种基于推理依据的剪枝策略,更好地对齐证据与决策。在眼动视频数据集上的基准测试表明,该方法不仅实现了最高的预测准确率,还在获得双重正确预测方面优于现有方法。本工作旨在推动高效且可靠的VLM研究,确保以准确性为导向的进展符合负责任人机交互与具身智能所需的透明性、可审计性和安全性要求。
原文摘要 · Abstract (English)
The rapid rise of Vision-Language Models (VLMs) in egocentric visual understanding has made low-latency inference in human-robot collaborative (HRC) tasks increasingly critical. Weight pruning techniques developed for VLMs to shrink model size and computation can be readily applied to satisfy the efficiency demands of on-board processing and real-time interactive robotics. Moreover, safe human-robot interaction demands pruning strategies that preserve doubly-correct predictions; outputs must be both accurate and evidentially grounded to mitigate risks and ensure user trust. In this paper, we present a new study of VLM pruning through the lens of doubly-correct prediction. Our experiments surprisingly show that existing pruning methods often preserve the right evidence localization but undermine correct prediction. To address this, we propose a rationale-informed pruning strategy that better aligns evidence with decisions. Benchmark results on egocentric video datasets demonstrate that our method not only achieves the highest prediction accuracy but also outperforms existing approaches in attaining doubly-correct predictions. We aim to stimulate research on efficient and reliable VLMs, ensuring accuracy-driven advances align with the transparency, auditability, and safety required for responsible human-robot interaction and embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。