压缩思维链会削弱模型安全性和抗幻觉能力,需同时优化效率与可信度。
Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression
- 对比多种压缩方法在推理过程中的可信度变化
- 压缩后模型在安全、幻觉抵抗等维度出现明显退化
- 提出新评估指标和对齐感知压缩方法,降低可信度损失
长思维链(Long-CoT)推理模型推动了推理轨迹压缩以降低推理开销,但现有评估主要关注任务准确率和令牌节省。信任属性(如安全、抗幻觉、多语言鲁棒性)与模型参数空间共存,压缩可能破坏这些属性而不限制准确率。我们首次系统研究压缩对模型可信度的影响,评估不同规模模型在三个维度的表现:安全性、幻觉抵抗、多语言鲁棒性。在受控比较中发现,压缩常导致可信度下降,且不同方法退化模式差异显著。为此提出归一化效率评分,揭示传统单一指标可能掩盖可信度权衡。作为验证,引入一种对齐感知的DPO变体,在推理基准上压缩率19.3%,可信度损失显著更小。结果表明,压缩应同时优化效率与可信度,二者同等重要。
原文摘要 · Abstract (English)
Long chain-of-thought (Long-CoT) reasoning models have motivated a growing body of work on compressing reasoning traces to reduce inference cost, yet existing evaluations focus almost exclusively on task accuracy and token savings. Trustworthiness properties, whether acquired or reinforced through post-training, are encoded in the same parameter space that compression modifies. This means preserving accuracy does not, a priori, guarantee preserving trustworthiness. We conduct the first systematic empirical study of how CoT compression affects model trustworthiness, evaluating multiple models of different scales along three dimensions: safety, hallucination resistance, and multilingual robustness. Under controlled comparisons, we find that CoT compression frequently introduces trustworthiness regressions and that different methods exhibit markedly different degradation profiles across dimensions. To enable fair comparison across bases, we propose a normalized efficiency score for each dimension that reveals how naïve scalar metrics can obscure trustworthiness trade-offs. As an existence proof, we further introduce an alignment-aware DPO variant that reduces CoT length by 19.3\% on reasoning benchmarks with substantially smaller trustworthiness loss. Our findings suggest that CoT compression should be optimized not only for efficiency but also for trustworthiness, treating both as equally important design constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。