arXiv:2412.11625cs.CL2024-12被引 6

用户更喜欢模型说假话但自信,而非承认不知

Fool Me, Fool Me: User Attitudes Toward LLM Falsehoods

论文配图:Fool Me, Fool Me: User Attitudes Toward LLM Falsehoods
图 1 · 摘自论文原文
  • 对比标记假话与不标记假话的响应偏好
  • 61%用户倾向未标注假话,69%偏好自信假话
  • 需判断真伪时偏好下降仍高于半数,适合关注人机交互者

尽管大语言模型在多个领域已成为核心工具,但常提供不准确或虚假信息。本研究考察用户对模型假话回应的偏好:比较明确标注假话与未标注假话的响应,以及自信假话与模型承认无知声明之间的偏好。此外,还探究要求用户评估陈述真实性如何影响这些偏好。结果显示,61%的用户更偏好未标注假话,69%更倾向自信假话。所有实验共300名用户参与,数据支持结论。当用户需评估真伪时,对未标注假话和假话的偏好略有下降但仍保持高位。这表明用户偏好可能通过反馈机制间接鼓励模型生成虚假内容,未来研究应关注其伦理与实践影响。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) have become central tools in various fields, they often provide inaccurate or false information. This study examines user preferences regarding falsehood responses from LLMs. Specifically, we evaluate preferences for LLM responses where false statements are explicitly marked versus unmarked responses and preferences for confident falsehoods compared to LLM disclaimers acknowledging a lack of knowledge. Additionally, we investigate how requiring users to assess the truthfulness of statements influences these preferences. Surprisingly, 61\% of users prefer unmarked falsehood responses over marked ones, and 69\% prefer confident falsehoods over LLMs admitting lack of knowledge. In all our experiments, a total of 300 users participated, contributing valuable data to our analysis and conclusions. When users are required to evaluate the truthfulness of statements, preferences for unmarked and falsehood responses decrease slightly but remain high. These findings suggest that user preferences, which influence LLM training via feedback mechanisms, may inadvertently encourage the generation of falsehoods. Future research should address the ethical and practical implications of aligning LLM behavior with such preferences.

大模型用户偏好假话生成人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。