首个基于真实佩戴数据的智能眼镜多模态问答基准,验证了现有模型在实际场景中的不足。
SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses

- 构建真实佩戴设备采集的多模态问答数据集,包含2422对图像与问题。
- 26个主流视觉语言模型在该基准上表现差距显著,平均准确率不足60%。
- 提出SUPERLENS智能代理,通过目标检测+查询拆分+多模态搜索提升性能。
AI驱动的智能眼镜快速发展,成为热门可穿戴设备,其核心应用之一是基于外部知识源的视觉问答(VQA)。现有适配智能眼镜的视觉语言模型(VLMs)通常在传统多模态数据集上训练与评估,但这些数据集缺乏真实使用场景的多样性与复杂性,且未体现关键挑战:必须先准确定位关注物体,才能进行外部知识检索。为此,我们提出SUPERGLASSES,首个完全由智能眼镜设备采集的真实世界数据构建的综合性VQA基准。SUPERGLASSES包含2,422个第一人称视角的图像-问题对,覆盖14个图像领域和8类查询,附带完整的搜索轨迹与推理标注。我们在该基准上评估了26个代表性VLMs,发现显著性能差距。为解决现有模型局限,我们进一步提出SUPERLENS——一种融合自动目标检测、查询解耦与多模态网络搜索的智能眼镜代理,实现增强型答案生成。SUPERLENS达到当前最优性能,优于GPT-4o 2.19%,凸显针对智能眼镜任务设计专用解决方案的重要性。数据集已公开于https://huggingface.co/datasets/xandery/SuperGlasses。
原文摘要 · Abstract (English)
The rapid advancement of AI-powered smart glasses-one of the hottest wearable devices-has unlocked new frontiers for multimodal interaction, with Visual Question Answering (VQA) over external knowledge sources emerging as a core application. Existing Vision Language Models (VLMs) adapted to smart glasses are typically trained and evaluated on traditional multimodal datasets; however, these datasets lack the variety and realism needed to reflect smart glasses usage scenarios and diverge from their specific challenges, where accurately identifying the object of interest must precede any external knowledge retrieval. To bridge this gap, we introduce SUPER- GLASSES, the first comprehensive VQA benchmark built on real-world data entirely collected by smart glasses devices. SUPERGLASSES comprises 2,422 egocentric image-question pairs spanning 14 image domains and 8 query categories, enriched with full search trajectories and reasoning annotations. We evaluate 26 representative VLMs on this benchmark, revealing significant performance gaps. To address the limitations of existing models, we further propose the SUPERLENS, a multimodal smart glasses agent that enables retrieval-augmented answer generation by integrating automatic object detection, query decoupling, and multimodal web search. SUPERLENS achieves state-of-the-art performance, outperforming GPT-4o by 2.19%, underscoring the need for task-specific solutions in smart glasses VQA. Our dataset is publicly available at https://huggingface.co/datasets/xandery/SuperGlasses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。