揭示大模型生成内容中的认知偏差及其对决策影响
Quantifying Cognitive Bias Induction in LLM-Generated Content
- 通过新构建数据集测试五类大模型的偏见诱导能力
- 26.42%内容改变原意情感,60.33%出现事实幻觉
- 适合关注AI决策风险与内容安全的研究者
大型语言模型被广泛应用于购物评价、摘要生成和医疗诊断辅助等场景,其输出会影响人类判断。本文研究大模型生成内容对用户认知偏见的影响,评估五类大模型在摘要和新闻事实核查任务中的表现,使用新构建的自更新数据集检验模型上下文一致性及幻觉倾向。结果表明,大模型在26.42%的情况下改变原始文本情感(框架偏见),在60.33%的超出知识截止时间的问题上产生幻觉,且在10.12%的情况下过度依赖提示词早期内容(优先效应)。此外,阅读大模型生成的评价摘要后,用户购买同一商品的可能性提高32%。为缓解该问题,本文评估了18种干预方法在三类模型上的效果,验证了针对性措施的有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) are integrated into applications like shopping reviews, summarization, or medical diagnosis support, where their use affects human decisions. We investigate the extent to which LLMs expose users to biased content and demonstrate its effect on human decision-making. We assess five LLM families in summarization and news fact-checking tasks, evaluating the consistency of LLMs with their context and their tendency to hallucinate on a new self-updating dataset. Our findings show that LLMs expose users to content that changes the context's sentiment in 26.42% of cases (framing bias), hallucinate on 60.33% of post-knowledge-cutoff questions, and highlight context from earlier parts of the prompt (primacy bias) in 10.12% of cases, averaged across all tested models. We further find that humans are 32% more likely to purchase the same product after reading a summary of the review generated by an LLM rather than the original review. To address these issues, we evaluate 18 mitigation methods across three LLM families and find the effectiveness of targeted interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。