arXiv:2508.18177cs.CVcs.LG2025-08被引 3

为视障者设计高效多智能体系统,让设备更懂环境、反应更快。

Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance

  • 用跨模态分层量化压缩大模型,内存从38GB降到11.3GB
  • 多智能体系统实现场景感知与记忆推理,语音响应延迟仅2.83-3.52秒
  • 在低内存下仍保持高精度,适合部署在移动设备上

视障人士在环境感知方面面临巨大挑战。传统辅助技术缺乏自适应智能,多聚焦于单一功能模块而非集成系统。尽管视觉语言模型(VLMs)为实现更丰富的整合理解提供了可能,但其部署受限于巨大的计算需求,需数十吉字节内存。为此,本文提出双技术创新框架:跨模态分层量化VLMs与场景感知的向量记忆多智能体系统。量化框架采用差异化策略,将内存从38GB降至11.3GB;多智能体系统利用向量记忆与感知-记忆-推理工作流,提供超出当前视野的环境信息,实现2.83-3.52秒的初始语音输出延迟。实验表明,量化后的190亿参数模型在MMBench上性能仅下降2.05%,在OCR-VQA上准确率达63.7%(原为64.9%),优于同等内存下的小模型。本研究显著提升了计算效率与辅助能力,覆盖场景感知、文字识别与导航等关键任务。

原文摘要 · Abstract (English)

Visually impaired individuals face significant challenges in environmental perception. Traditional assistive technologies often lack adaptive intelligence, focusing on individual components rather than integrated systems. While Vision-Language Models (VLMs) offer a promising path to richer, integrated understanding, their deployment is severely limited by substantial computational requirements, demanding dozens of gigabytes of memory. To address these gaps in computational efficiency and integrated design, this study proposes a dual technological innovation framework: a cross-modal differentiated quantization framework for VLMs and a scene-aware vectorized memory multi-agent system. The quantization framework implements differentiated strategies, reducing memory from 38GB to 11.3GB. The multi-agent system uses vectorized memory and perception-memory-reasoning workflows to provide environmental information beyond the current view, achieving 2.83-3.52s latency to initial speech output. Experiments show the quantized 19B-parameter model only experiences a 2.05% performance drop on MMBench and maintains 63.7 accuracy on OCR-VQA (original: 64.9), outperforming smaller models with equivalent memory. This research advances computational efficiency and assistive technology, offering comprehensive assistance in scene perception, text recognition, and navigation.

视障辅助多智能体量化VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。