首个面向可穿戴设备的视觉问答基准,测试真实场景下AI在模糊、遮挡等挑战下的表现。
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios
- 构建真实佩戴场景的图像-问题-答案三元组,模拟日常使用中的视觉难题。
- 多模态模型在该基准上准确率仅24-52%,尤其在低质量图像和推理任务中大幅下降。
- 适合研究可穿戴设备、边缘AI与鲁棒视觉理解的开发者与学者使用。
我们提出WearVQA,首个专为智能眼镜等可穿戴设备上的多模态AI助手设计的视觉问答评估基准。不同于以往聚焦高质量第三人称图像的评测,WearVQA反映第一人称交互的独特挑战——视觉输入可能被遮挡、光照差、未缩放或模糊,且问题源于真实的可穿戴使用场景。该基准包含2,520个精心筛选的图像-问题-答案三元组,覆盖7个不同图像领域(包括以文本为中心和通用场景)、10类认知任务(从基础识别到多种推理形式),以及6种常见的可穿戴设备图像质量问题。所有问题均仅依赖视觉输入与常识即可回答。配套的LLM-as-a-judge评估框架达到96%标注准确率。开源与专有多模态大模型在WearVQA上表现不佳,准确率仅为24-52%,在低质量图像和高阶推理任务中下降明显。这些结果使WearVQA成为推动真实世界多模态可穿戴系统技术进步的重要挑战性基准。
原文摘要 · Abstract (English)
We introduce WearVQA, the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-model AI assistant on wearable devices like smart glasses. Unlike prior benchmarks that focus on high-quality, third-person imagery, WearVQA reflects the unique challenges of ego-centric interaction-where visual inputs may be occluded, poorly lit, unzoomed, or blurry, and questions are grounded in realistic wearable use cases. The benchmark comprises 2,520 carefully curated image-question-answer triplets, spanning 7 diverse image domains including both text-centric and general scenes, 10 cognitive task types ranging from basic recognition to various forms of reasoning, and 6 common wearables-specific image quality issues. All questions are designed to be answerable using only the visual input and common senses. WearVQA is paired with a rigorous LLM-as-a-judge evaluation framework with 96% labeling accuracy. Open-source and proprietary multi-model LLMs achieved a QA accuracy as low as 24-52% on WearVQA, with substantial drops on lower-quality images and reasoning-heavy tasks. These observations position WearVQA as a comprehensive and challenging benchmark for guiding technical advancement towards robust, real-world multi-model wearables AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。