arXiv:2609.04720cs.CLcs.AI2026-09

提出新基准KoNA,评估视觉语言模型对部分错误问题的精准拒绝能力。

Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models

论文配图:Knowing What Not to Answer: Selective Non-Compliance in Vision-Language Models
图 1 · 摘自论文原文
  • 设计五类混合问题,测试模型能否只拒绝不可答部分
  • 多模型实测显示90%以上在复杂问题上误回答
  • 微调后模型在保持准确率同时显著提升拒绝精度

视觉语言模型应仅对合理请求响应,对错误、危险或无法回答的问题拒绝。但现有评测将请求整体视为可答或不可答,忽略真实问题中常混有可答与不可答成分。本文提出KoNA基准,覆盖五类场景:虚假前提、视觉不可及、普遍未知、任务可行性与安全风险。通过成对单成分与复合问题,评估模型在查询级与组件级的拒绝能力。跨多种模型的实验表明,多数模型在需要选择性拒绝时表现不佳,错误率超90%。为此,我们用KoNA中的选择性拒绝样本与完全可答样本联合微调模型,结果在保持原有问答性能基础上,非合规准确率显著提升,证明模型能区分可答与不可答成分并做出恰当回应。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are expected to respond helpfully to appropriate requests while withholding compliance with requests that are incorrect, unsafe, infeasible, or unanswerable. However, existing benchmarks predominantly evaluate non-compliance at the level of the query as a whole, assuming that each request either warrants compliance or requires withholding compliance. In practice, real-world queries can contain a mixture of answerable content and components for which compliance should be withheld. In this paper, we introduce KoNA, a benchmark for evaluating selective non-compliance in VLMs across five categories: False Premise, Visual Inaccessibility, Universal Unknown, Task Feasibility, and Safety. Each task evaluates two capabilities: query-level non-compliance and component-level non-compliance under paired single and compound queries. Our evaluation across diverse VLMs shows that models often fail to refuse, correct, or abstain appropriately, and these failures become more pronounced when queries require selective non-compliance. To address this challenge, we fine-tune VLMs using KoNA examples that require selective non-compliance, together with a fully answerable set that should receive direct answers. Our fine-tuned models achieve substantial improvements in non-compliance accuracy while largely maintaining performance on fully answerable tasks. These results suggest that the fine-tuned models can distinguish between answerable components and those requiring non-compliance and respond in a task-appropriate manner.

视觉语言模型拒绝机制评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。