arXiv:2512.17920cs.CLcs.AI2025-12被引 1

提出新基准,分离评估大模型在压缩下的指令遵循与语义准确度。

Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression

  • 设计独立测量约束遵循与语义准确度的测试框架
  • 发现中等压缩下约束违反率最高(97.2%),极端压缩反而表现更好
  • 揭示强化学习对齐导致指令偏离,适合模型部署优化研究者

大语言模型在提示压缩下性能下降,但机制尚不明确。本文提出压缩衰减理解测试(CDCT),在5种压缩等级(从极强c=0.0,约2词到无压缩c=1.0,约135词)下,独立评估9个前沿模型在8个概念上的约束遵循(CC)与语义准确度(SA)。三名大模型评委对约束遵循达成近乎完美的评分一致性(Fleiss' κ=0.90)。观察到约束遵循呈现普遍的U型曲线模式(97.2%出现),在中等压缩(c=0.5,约27词)时违反率最高。反直觉的是,模型在极端压缩下表现优于中等长度。两维度统计正交(r=0.193, p=0.084),约束影响是语义影响的2.9倍。通过强化学习人类反馈(RLHF)消融实验验证:移除‘帮助性’信号后,约束遵循平均提升598%(71/72次试验,p<0.001),79%达到完全遵循。这表明,强化学习对齐中的帮助性行为是中等压缩下约束违背的主要原因。推理模型比高效模型表现高27.5%(Cohen's d=0.96)。研究揭示了强化学习对齐与指令遵循间的根本矛盾,为实际系统优化提供可操作指导。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit degraded performance under prompt compression, but the mechanisms remain poorly understood. We introduce the Compression-Decay Comprehension Test (CDCT), a benchmark that independently measures constraint compliance (CC) and semantic accuracy (SA) across compression levels. We evaluate 9 frontier LLMs across 8 concepts using 5 compression levels from extreme (c=0.0, ~2 words) to none (c=1.0, ~135 words). A three-judge LLM jury achieves almost perfect inter-rater agreement on CC (Fleiss' \k{appa}=0.90). We observe a universal U-curve pattern in constraint compliance (97.2% prevalence), with violations peaking at medium compression (c=0.5, ~27 words). Counterintuitively, models perform better at extreme compression than medium lengths. The dimensions are statistically orthogonal (r=0.193, p=0.084), with constraint effects 2.9x larger than semantic effects. Experimental validation via RLHF ablation confirms our constraint salience hypothesis: removing "helpfulness" signals improves CC by 598% on average (71/72 trials, p<0.001), with 79% achieving perfect compliance. This demonstrates that RLHF-trained helpfulness behaviors are the dominant cause of constraint violations at medium compression. Reasoning models outperform efficient models by 27.5% (Cohen's d=0.96). Our findings reveal a fundamental tension between RLHF alignment and instruction-following, providing actionable guidelines for improving deployed systems.

大模型评测指令遵循压缩感知强化学习对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。