小模型虽能答对题,却常无视反向指令,说明答题能力不等于听从指令。
Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models
- 设计冲突指令测试模型是否遵从用户要求
- 小模型准确率高但常忽略反向指令,大模型差距更明显
- 建议评估时加入指令遵循率,避免掩盖行为缺陷
指令微调旨在让语言模型遵循用户请求,但小模型在指令与常规任务行为冲突时是否仍能遵守尚不明确。本文在多项选择题问答(MCQA)、情感分类和数学问题解答三个任务上,分别设置标准指令与冲突的非标准指令(如选择错误选项、输出相反情感、返回两倍答案)。通过跨任务设计,检验模型抗拒冲突指令的行为是否与特定任务相关,或体现普遍倾向。所有预测均以原始真实标签评分,若模型忽略非标准指令,仍可能显示高准确率。采用标准准确率、非标准准确率及指令遵循失败率(IFFR)评估不同规模的Qwen模型。结果显示,标准准确率与指令遵循能力随模型规模提升,但并非所有任务和数据集都一致。小模型保持任务胜任力,但频繁忽略非标准指令;大模型在两种情境间差距显著。这表明任务能力提升并不自动带来可靠的行为控制。任务胜任力与指令遵循是两种独立能力,仅报告标准准确率会掩盖指令遵循失败。
原文摘要 · Abstract (English)
Instruction tuning is meant to make language models follow user requests, yet it is unclear whether small models comply when an instruction conflicts with their usual task behavior. We study this across three tasks - multiple-choice question answering (MCQA), sentiment classification, and mathematical question answering - by pairing a standard instruction with a conflicting non-standard one (select an incorrect option, output the opposite sentiment, or return twice the answer). This cross-task design allows us to test whether resistance to conflicting instructions is tied to specific task characteristics or reflects a broader behavioral tendency. As all predictions are scored against the original ground truth, a model that ignores the non-standard instruction still appears accurate. Using standard accuracy, non-standard accuracy, and an Instruction-Following Failure Rate (IFFR), we evaluate instruction-tuned Qwen models across sizes. Both standard accuracy and instruction following generally improve with scale, although the pattern is not consistent across all tasks and datasets. Small models stay competent yet routinely ignore the non-standard instruction, while larger models show a clear gap between the two settings. These findings suggest that gains in task capability do not automatically provide reliable control over model behavior. Task competence and instruction following are therefore distinct abilities, and reporting only standard accuracy hides instruction-following failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。