arXiv:2608.14630cs.CLcs.AI2026-08

语言模型的表达方式可能诱导医生误判,即使内容正确也不安全。

Characterizing Rhetorical Misalignment in Decision-Making with Language Models

论文配图:Characterizing Rhetorical Misalignment in Decision-Making with Language Models
图 1 · 摘自论文原文
  • 构建决策理论框架,分析语言模型在决策中不当修辞的影响
  • 实验证明不同模型平均导致2.81%的错误决策翻转
  • 适合关注AI医疗安全、认知偏见与提示工程的研究者

人类决策常受多种已知认知偏见影响。随着大语言模型(LLMs)越来越多地融入高风险人机决策场景,需明确其输出是否会放大潜在偏见、如何影响人类判断,以及是否引发有害后果。本文提出一种决策理论框架,研究「修辞错位」——即模型在特定决策情境下使用不恰当的语言表达,从而诱导次优人类决策的现象。通过基于美国医学执照考试数据集的真实临床决策实验,我们发现不同模型平均导致2.81%的决策翻转,即医生从正确答案转向错误答案。参与者报告的推理表明,这些改变与模型语言诱发的锚定效应、权威偏见和损失厌恶等认知偏见密切相关。为实现可扩展评估,我们使用由大语言模型模拟的决策者,计算量化修辞错位程度。研究揭示了高风险领域中此前未被重视的安全隐患:模型虽事实正确,仍可能因表达方式造成伤害。

原文摘要 · Abstract (English)

Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether their outputs can amplify potential biases, how this influences human decisions, and crucially, whether it can lead to harmful consequences. In this work, we develop a decision-theoretic framework to study rhetorical misalignment, a failure mode where an LLM uses rhetorically inappropriate forms of presentation for a given decision context, thereby inducing suboptimal human decisions. We empirically investigate this phenomenon through a human-subject experiment in realistic clinical decision-making using a dataset curated from the United States Medical Licensing Examination. By measuring how LLM-generated information affects decisions, we observe that LLMs induce an average 2.81% rate of harmful decision flips across different models, where clinician participants change from a correct to an incorrect answer. Rationales reported by participants provide evidence that these revisions are closely related to the language used by LLMs that may induce different types of cognitive biases, including anchoring, authority bias, and loss aversion. To enable scalable evaluation, we instantiate our theoretical framework using decision-makers simulated by LLMs to computationally measure rhetorical misalignment. Our findings reveal a safety concern previously unrecognized in high-stakes domains: a model can be factually aligned yet still induce harm through its rhetorical presentation.

AI安全认知偏见医疗AI修辞错位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。