arXiv:2609.03331cs.CL2026-09

测试视觉语言模型在持续错误前提下的纠错与协作能力

FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models

论文配图:FPCO-Dialog: A Multi-Turn False-Premise Benchmark for Correction and Cooperation in Vision-Language Models
图 1 · 摘自论文原文
  • 构建多轮对话基准,固定错误前提持续出现
  • 20个模型表现差异显著,纠错率受前提类型影响
  • 适合评估模型在真实对话中的可靠性与鲁棒性

视觉语言模型越来越多地应用于多轮对话场景,用户可能基于错误假设描述视觉内容。然而现有评估很少分离出相同视觉错误前提在多轮中持续出现时模型的响应行为。本文提出FPCO-Dialog,一个用于评估视觉语言模型在重复错误前提下纠错与协作能力的基准。该数据集包含1,080张图像和10,800个问题轮次,按视觉复杂度、物体类别和错误前提类型分层,并采用10轮对话协议:先给出正确对话前缀,随后反复使用错误前提指代表达。我们采用无模型依赖的评估协议和CorrTP@K指标(由两个独立检测器评分),对20个商用及开源视觉语言模型进行评测。结果揭示了模型间在整体纠错倾向、逐轮动态行为以及不同错误前提类型下的系统性差异。数据集、评估协议、模型输出、检测标签和代码均已公开。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly deployed in multi-turn settings where users may describe visual content with incorrect assumptions. Yet existing evaluations rarely isolate how models respond when the same visually grounded false premise persists across dialogue turns. We introduce FPCO-Dialog, a benchmark for evaluating correction and cooperation behavior in VLMs under repeated false premises. FPCO-Dialog contains 1,080 images and 10,800 question turns, stratified by visual complexity, object category, and false-premise class, and uses a 10-turn protocol in which a correct dialogue prefix is followed by repeated false-premise referring expressions. We evaluate 20 commercial and open-source VLMs with a model-agnostic protocol and CorrTP@K, a correction-rate metric over false-premise turns, scored by two independent detectors. FPCO-Dialog reveals substantial and persistent cross-model differences in aggregate correction tendency, model-specific turn-wise dynamics, and systematic variation across false-premise types under the benchmark's substitution distribution. The dataset, evaluation protocol, model outputs, detector labels, and code are available.

视觉语言模型多轮对话纠错能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。