arXiv:2606.07951cs.CLcs.AI2026-06

语言模型重写时会无意识夸大语气,影响可信度。

From `May' to `Is': Certainty Distortion in Language Model Rewriting

论文配图:From `May' to `Is': Certainty Distortion in Language Model Rewriting
图 1 · 摘自论文原文
  • 用新指标评估模型语气变化,贴近人类判断。
  • 75%输出出现语气失真,多数情况语气被放大1.5到2倍。
  • 反复改写会让语气越来越强,医疗领域尤为明显。

随着人类越来越多地依赖语言模型进行科学、新闻和医疗信息的讨论、重写与摘要,其表达信心的准确性变得至关重要。本文研究了语言模型在保持语义不变的前提下,表达确定性发生显著变化的现象——即确信度扭曲。我们提出一种基于语言模型的评估指标,该指标与群体层面的人类确信判断一致。通过该指标,我们分析了不同规模和架构的语言模型在科学与医学传播任务中的确信度扭曲现象。结果表明,高达75%的模型输出存在确信度扭曲,且在重写任务中系统性地偏向增强语气:多数模型将确信度提高的概率是降低的1.5至2倍。这种效应在多次重写中会累积:在医疗领域,Claude Haiku 4-5在单次重写后使20%的陈述确信度提升,五次迭代后上升至40%。基于提示的干预虽能减少扭曲,但无法根除。这些发现揭示了语言模型普遍存在夸大语气的倾向,对高风险场景下的使用者具有重要警示意义。

原文摘要 · Abstract (English)

Humans increasingly turn to Language Models (LMs) in ways that shape beliefs and drive decisions, including discussing, rewriting, and summarizing information from scientific articles, news, and medical reports. However, in these domains, where how confidently a claim is expressed matters, little is known about whether LMs faithfully preserve it. In this work, we investigate certainty distortion in LMs, defined as meaningful changes in expressed certainty when semantic content is preserved. We propose an LM-based evaluation metric that is consistent with population-level judgments of certainty. Using this metric, we characterize certainty distortion across different sizes and families of models in the context of scientific and medical communication tasks. Our results show that certainty distortion affects up to 75\% of LM outputs and is systematically asymmetric in rewriting tasks with most LMs being 1.5-2$\times$ more likely to increase the expressed certainty than to decrease it. These effects can compound over repeated paraphrasing: in the medical domain, claude-haiku-4-5 increases certainty of 20\% examples after a single iteration, increasing to 40\% after five iterations. Prompt-based interventions reduce overall certainty distortion but do not eliminate it. Together, these findings reveal a general bias toward inflating expressed certainty, with direct implications for users who rely on LMs in high-stakes domains.

语言模型确信度重写偏差医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。