arXiv:2602.13479cs.CVcs.HC2026-02

让可穿戴设备实时识别文本并理解上下文,功耗仅需全分辨率的49%

GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables

  • 分层处理:关键文本用高分辨率识别,背景用低分辨率传流
  • 在5类任务上达72%准确率,功耗降低至49%
  • 适合资源受限的可穿戴设备实时视觉问答场景

视频大语言模型在文本识别和基于文本的视觉问答(Text VQA)任务中表现优异。然而,在可穿戴设备上部署Text VQA面临根本矛盾:文本识别需要高分辨率视频,但高清视频流会快速耗尽电池并引发过热降频。此外,现有模型难以在实时视频流中保持跨帧文本的连贯上下文。我们观察到文本识别与视觉推理对分辨率的需求不对称——OCR需要精细细节,而场景理解可容忍粗粒度特征。为此,我们提出一种混合架构:本地选择性地进行高分辨率OCR,同时以低分辨率视频流传输视觉上下文。在涵盖五类任务的Text VQA基准测试中,系统实现72%准确率,功耗仅为全分辨率流的0.49倍,使资源受限的可穿戴设备得以持续运行VQA任务,且不牺牲文本理解质量。

原文摘要 · Abstract (English)

Video Large Language Models (Video LLMs) have shown remarkable progress in understanding and reasoning about visual content, particularly in tasks involving text recognition and text-based visual question answering (Text VQA). However, deploying Text VQA on wearable devices faces a fundamental tension: text recognition requires high-resolution video, but streaming high-quality video drains battery and causes thermal throttling. Moreover, existing models struggle to maintain coherent temporal context when processing text across multiple frames in real-time streams. We observe that text recognition and visual reasoning have asymmetric resolution requirements - OCR needs fine detail while scene understanding tolerates coarse features. We exploit this asymmetry with a hybrid architecture that performs selective high-resolution OCR on-device while streaming low-resolution video for visual context. On a benchmark of text-based VQA samples across five task categories, our system achieves 72% accuracy at 0.49x the power consumption of full-resolution streaming, enabling sustained VQA sessions on resource-constrained wearables without sacrificing text understanding quality.

可穿戴计算文本识别视觉问答能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。