用颜色标记每个短语的真实性,让大模型输出更透明可信。
Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
- 将回复中每个短语按真实度着色,直观展示信息可信度。
- 208人实验显示,着色方案显著提升用户信任与验证效率。
- 适合需要增强模型可解释性的应用开发者参考使用。
大型语言模型(LLMs)容易生成不准确或虚假信息,即所谓的“幻觉”或“虚构”。尽管已有技术进步用于评估模型回答的事实性,但关于如何有效向用户传达这些信息的研究仍有限。为此,我们开展了两项基于场景的实验,共208名参与者,系统比较了多种传达事实性评分的设计策略,评估用户对信任度、验证准确性难易度及偏好度的反馈。结果表明,所有短语均根据事实性评分进行颜色编码的设计最受青睐,用户更信任该方案,并且更容易验证回答的准确性。本研究为大模型应用开发者和设计师提供了实用的设计指南,旨在校准用户信任、契合用户偏好,并提升用户对模型输出的审查能力。
原文摘要 · Abstract (English)
Large language models (LLMs) are susceptible to generating inaccurate or false information, often referred to as "hallucinations" or "confabulations." While several technical advancements have been made to detect hallucinated content by assessing the factuality of the model's responses, there is still limited research on how to effectively communicate this information to users. To address this gap, we conducted two scenario-based experiments with a total of 208 participants to systematically compare the effects of various design strategies for communicating factuality scores by assessing participants' ratings of trust, ease in validating response accuracy, and preference. Our findings reveal that participants preferred and trusted a design in which all phrases within a response were color-coded based on factuality scores. Participants also found it easier to validate accuracy of the response in this style compared to a baseline with no style applied. Our study offers practical design guidelines for LLM application developers and designers, aimed at calibrating user trust, aligning with user preferences, and enhancing users' ability to scrutinize LLM outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。