用图像识别任务测试大模型对家居场景的感知与人类是否一致。
Vision Language Models as Values Detectors
- 用12张家居图让14人标注关键元素,对比5个大模型输出。
- 最强模型LLaVA 34B与人类一致率仍较低,但能识别有情感价值的细节。
- 适合做社会机器人、辅助设备的感知系统优化参考。
融合文本与视觉输入的大语言模型为复杂数据解读带来了新可能。尽管这些模型基于视觉刺激生成连贯且上下文相关文本的能力出色,但其在识别图像中相关元素时与人类感知的对齐程度仍需深入探究。本文研究了先进大模型与人类标注者在家庭环境场景中识别相关元素的一致性。我们创建了12张描绘不同居家场景的图像,并邀请14名标注者识别每张图中的关键元素。随后将人类结果与5种大模型(包括GPT-4o和四种LLaVA变体)的输出进行对比。结果显示,模型与人类的对齐程度参差不齐,其中LLaVA 34B表现最优但一致性依然偏低。然而,结果分析表明,模型具备检测图像中蕴含价值要素的潜力,若通过改进训练和优化提示,有望在社会机器人、辅助技术及人机交互中提供更深层洞察和更贴合语境的响应。
原文摘要 · Abstract (English)
Large Language Models integrating textual and visual inputs have introduced new possibilities for interpreting complex data. Despite their remarkable ability to generate coherent and contextually relevant text based on visual stimuli, the alignment of these models with human perception in identifying relevant elements in images requires further exploration. This paper investigates the alignment between state-of-the-art LLMs and human annotators in detecting elements of relevance within home environment scenarios. We created a set of twelve images depicting various domestic scenarios and enlisted fourteen annotators to identify the key element in each image. We then compared these human responses with outputs from five different LLMs, including GPT-4o and four LLaVA variants. Our findings reveal a varied degree of alignment, with LLaVA 34B showing the highest performance but still scoring low. However, an analysis of the results highlights the models' potential to detect value-laden elements in images, suggesting that with improved training and refined prompts, LLMs could enhance applications in social robotics, assistive technologies, and human-computer interaction by providing deeper insights and more contextually relevant responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。