评测大模型在助人场景中的视觉语言能力,发现其有潜力但仍有局限。
Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

- 用头戴相机采集真实场景数据,构建助人应用基准测试
- 模型在识币、读文字、跨语言理解任务中表现不一,部分任务准确率不足60%
- 适合研究无障碍技术、多模态系统落地的开发者参考
多模态大语言模型(MLLMs)通过融合视觉编码器与大规模语言模型,显著提升了图像理解能力,在图像描述、视觉问答和多模态对话等任务上表现出色,常可在零样本或少样本条件下运行。其通用性强、接口灵活,被视为现实世界视觉-语言应用的理想基础。辅助人工智能旨在通过自然语言帮助用户与环境交互,这需要强大的视觉识别、上下文推理和多语言理解能力——这些正是MLLMs被寄予厚望的能力。然而,其在辅助场景中的实际效果尚不明确。本文通过评估先进MLLMs在真实任务中的表现,探索其是否可支撑辅助智能:包括识别日常物品(如货币)、基于场景文本回答问题、跨语言阅读视觉内容。为此,我们开发了名为NetraLink的系统,利用头戴式GoPro捕捉真实世界的视角数据,并构建了一个涵盖上述场景的基准数据集。研究结果全面诊断了当前MLLMs的表现,揭示了其在基于视觉感知与语言交互的辅助技术中的优势与短板。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong performance on tasks like image captioning, visual question answering, and multimodal dialogue, often in zero- and few-shot settings. Their general-purpose capabilities and flexible interfaces make MLLMs a promising foundation for real-world vision-language applications. Assistive AI aims to help users interact with their environments through natural language. These scenarios demand robust visual recognition, contextual reasoning, and multilingual comprehension-capabilities that MLLMs are believed to offer. However, their effectiveness in assistive settings remains to be fully understood. In this work, we explore whether MLLMs can support Assistive AI by evaluating state-of-the-art models on real-world tasks: recognizing everyday objects like currency, answering questions based on scene text, and reading visually presented content across multiple languages. To this end, we developed a system, NetraLink, using a head-mounted GoPro to capture real-world egocentric data, and collected a benchmark covering these assistive scenarios. Our findings provide a comprehensive diagnostic of current MLLMs, highlighting their strengths and limitations in enabling assistive technologies grounded in visual perception and language interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。