研究大模型自解释的长短平衡,发现简短说明也能准确答题。
The Sufficiency-Conciseness Trade-off in LLM Self-Explanation from an Information Bottleneck Perspective
- 从信息瓶颈视角,将解释看作压缩后的关键信息
- 实验表明简短解释仍能保持高准确率,压缩过度才导致性能下降
- 覆盖英、波斯双语,验证结论跨语言有效性
大型语言模型越来越多地依赖自解释(如思维链推理)来提升多步问答的性能。尽管这些解释提高了准确性,但通常冗长且生成成本高,引发疑问:究竟需要多少解释才是必要的?本文从信息瓶颈角度出发,将解释视为保留正确答案所需关键信息的压缩表征。为此,我们构建了约束解释长度并利用多个语言模型在ARC Challenge数据集上评估充分性的评测流程。为拓展适用范围,实验同时在英语原版数据集和通过翻译得到的波斯语资源受限语言上进行。结果表明,更简洁的解释往往依然充分,在显著缩短长度的同时保持高准确率;而过度压缩则会导致性能下降。
原文摘要 · Abstract (English)
Large Language Models increasingly rely on self-explanations, such as chain of thought reasoning, to improve performance on multi step question answering. While these explanations enhance accuracy, they are often verbose and costly to generate, raising the question of how much explanation is truly necessary. In this paper, we examine the trade-off between sufficiency, defined as the ability of an explanation to justify the correct answer, and conciseness, defined as the reduction in explanation length. Building on the information bottleneck principle, we conceptualize explanations as compressed representations that retain only the information essential for producing correct answers.To operationalize this view, we introduce an evaluation pipeline that constrains explanation length and assesses sufficiency using multiple language models on the ARC Challenge dataset. To broaden the scope, we conduct experiments in both English, using the original dataset, and Persian, as a resource-limited language through translation. Our experiments show that more concise explanations often remain sufficient, preserving accuracy while substantially reducing explanation length, whereas excessive compression leads to performance degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。