测试大模型能否摆脱训练习惯,听懂反常规指令。
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
- 设计对抗性指令基准,检验模型是否能突破固有模式。
- 在23个领域构建1012条中英文高质量问题,验证模型适应性。
- 适合关注模型真实指令遵循能力的研究者与开发者。
大型语言模型在多种任务上表现优异,但常因监督微调阶段学习到的标准模式而表现出认知惯性,难以遵循与之冲突的指令。为此,我们提出Inverse IFEval基准,评估模型克服训练诱导偏见、服从对抗性指令的能力。该基准包含八类挑战,如问题修正、故意文本缺陷、无注释代码和反事实回答等。通过人机协作流程,我们在23个领域构建了1012条高质量中英文问题数据集,并采用优化的LLM-as-a-Judge框架进行评估。对现有主流大模型的实验表明,Inverse IFEval基准具有必要性。研究强调,未来的对齐工作不仅应追求流畅性和事实正确性,还需考虑模型在非常规情境下的适应能力。我们希望Inverse IFEval成为诊断工具,并为缓解认知惯性、减少对狭窄模式的过拟合、提升模型在多样化且不可预测的真实场景中的指令遵循可靠性提供基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) achieve strong performance on diverse tasks but often exhibit cognitive inertia, struggling to follow instructions that conflict with the standardized patterns learned during supervised fine-tuning (SFT). To evaluate this limitation, we propose Inverse IFEval, a benchmark that measures models Counter-intuitive Abilitytheir capacity to override training-induced biases and comply with adversarial instructions. Inverse IFEval introduces eight types of such challenges, including Question Correction, Intentional Textual Flaws, Code without Comments, and Counterfactual Answering. Using a human-in-the-loop pipeline, we construct a dataset of 1012 high-quality Chinese and English questions across 23 domains, evaluated under an optimized LLM-as-a-Judge framework. Experiments on existing leading LLMs demonstrate the necessity of our proposed Inverse IFEval benchmark. Our findings emphasize that future alignment efforts should not only pursue fluency and factual correctness but also account for adaptability under unconventional contexts. We hope that Inverse IFEval serves as both a diagnostic tool and a foundation for developing methods that mitigate cognitive inertia, reduce overfitting to narrow patterns, and ultimately enhance the instruction-following reliability of LLMs in diverse and unpredictable real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。