arXiv:2603.07868cs.AIcs.LG2026-03Conference of the …

评估视觉语言模型在酒店决策中的信息有用性,发现需微调才能有效理解关键视觉信息。

Hospitality-VQA: Decision-Oriented Informativeness Evaluation for Vision-Language Models

  • 提出'信息度'框架量化图像-问题对的决策相关价值
  • 构建专用酒店设施VQA数据集,覆盖多种设施类型
  • 揭示现有模型缺乏决策意识,微调后才具备可靠信息推理能力

视觉语言模型(VLMs)在通用领域表现出色,但在酒店等决策导向场景的应用仍不明确。本文研究了VLMs在酒店及设施图像上的视觉问答能力,这些图像直接影响消费者决策。现有VQA基准多关注事实正确性,却忽视用户实际所需信息。为此,我们提出'信息度'框架,用于量化图像-问题对提供的与酒店相关的有用信息量。基于此框架,构建了一个涵盖多种设施类型的专用酒店VQA数据集,问题设计贴合用户核心信息需求。在该基准上测试多个先进VLMs,发现模型本质缺乏决策意识,关键视觉信号未被充分利用;仅经少量领域特定微调后,信息推理能力才显著提升。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have demonstrated impressive multimodal understanding in general domains. However, their applicability to decision-oriented domains such as hospitality remains largely unexplored. In this work, we investigate how well VLMs can perform visual question answering (VQA) about hotel and facility images that are central to consumer decision-making. While many existing VQA benchmarks focus on factual correctness, they rarely capture what information users actually find useful. To address this, we first introduce Informativeness as a formal framework to quantify how much hospitality-relevant information an image-question pair provides. Guided by this framework, we construct a new hospitality-specific VQA dataset that covers various facility types, where questions are specifically designed to reflect key user information needs. Using this benchmark, we conduct experiments with several state-of-the-art VLMs, revealing that VLMs are not intrinsically decision-aware-key visual signals remain underutilized, and reliable informativeness reasoning emerges only after modest domain-specific finetuning.

视觉问答决策智能多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。