arXiv:2606.26101cs.CLcs.AI2026-06中稿 · as a regular paper…

构建多区域评测基准,区分模型回答与猜答,避免数据污染干扰。

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

论文配图:Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models
图 1 · 摘自论文原文
  • 设计五域1200项测试,标注答案预期与污染风险,分离回答与拒答行为。
  • 强指令微调模型仍无法完全实现从回答到拒答的可靠过渡,校准能力差。
  • 适用于评估大模型可靠性,尤其关注拒答合理性与数据污染影响。

可靠的大型语言模型评估需将有据回答与无据猜测区分开,避免数据污染、提示特异性或泛化拒答行为的干扰。本文提出一种抗污染的多区域评测基准,用于衡量在冻结训练时标签条件下,模型从可回答知识向预期拒答未知状态的转变过程。该基准涵盖五个领域共1200个样本,包含明确的拒答预期、污染风险元数据,并采用官方严格解析器与标准化鲁棒解析器双重解析。我们在锁定回答或拒答提示、仅回答控制及提示模板变体下评估FLAN-T5、Qwen2.5-Instruct和Llama-3-Instruct模型。该基准不会被通用不回答行为解决:FLAN基线在主动拒答方面表现弱,更强的指令微调模型虽呈现选择性但不完整的过渡。Qwen2.5-3B-Instruct整体表现最佳,但预期回答区域仍困难,校准效果不佳,且良性项目仍存在拒绝现象。提示与解析器鲁棒性分析保持主排名与定性结论不变。该基准提供了一套可复现的协议,用于审计回答能力、拒答行为、拒绝倾向与污染风险这四个相互关联的可靠性维度。数据集已公开于https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark。

原文摘要 · Abstract (English)

Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.

大模型评测拒答机制数据污染可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。