真实用户提问常不完整,模型理解能力被高估。
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
- 构建653个真实韩国网络提问数据集,含原始与改写版本
- 顶尖模型在原始提问上准确率不足50%,改写后提升8-22个百分点
- 小模型从明确化提问中获益更大,现有检索无法弥补信息缺失
当前视觉语言模型评测多基于结构清晰的明确问题,但真实用户提问常非正式且信息不足。用户依赖图像传递上下文,许多内容未明说。我们构建了HAERAE-Vision数据集,包含从8.6万候选中筛选出的653个韩国在线社区真实视觉问题,每个问题配有显式改写版,共生成1,306个查询变体。评估39个VLM发现,即使最先进的模型(GPT-5、Gemini 2.5 Pro)在原始问题上的准确率也低于50%。仅通过问题显式化,性能提升达8至22个百分点,小模型受益更显著。进一步实验表明,即便使用网络搜索,未明确问题的表现仍低于无搜索的明确问题,说明当前检索机制无法弥补用户未言明的信息。研究揭示,大量模型表现不佳源于自然提问的不完整,而非模型能力不足,暴露出评测与实际应用间的重大鸿沟。
原文摘要 · Abstract (English)
Current vision-language benchmarks predominantly feature well-structured questions with clear, explicit prompts. However, real user queries are often informal and underspecified. Users naturally leave much unsaid, relying on images to convey context. We introduce HAERAE-Vision, a benchmark of 653 real-world visual questions from Korean online communities (0.76% survival from 86K candidates), each paired with an explicit rewrite, yielding 1,306 query variants in total. Evaluating 39 VLMs, we find that even state-of-the-art models (GPT-5, Gemini 2.5 Pro) achieve under 50% on the original queries. Crucially, query explicitation alone yields 8 to 22 point improvements, with smaller models benefiting most. We further show that even with web search, under-specified queries underperform explicit queries without search, revealing that current retrieval cannot compensate for what users leave unsaid. Our findings demonstrate that a substantial portion of VLM difficulty stem from natural query under-specification instead of model capability, highlighting a critical gap between benchmark evaluation and real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。