研究大模型如何信任人类,发现其信任模式与人类相似但存在偏见。
A closer look at how large language models trust humans: patterns and biases
- 基于行为理论分析大模型对人类能力、善意和诚信的信任机制。
- 43,200次模拟实验显示,信任主要受可信度影响,但年龄、性别等有偏见。
- 新模型在金融场景中更易受人口特征影响,需警惕潜在风险。
随着大语言模型(LLMs)及其代理在决策场景中越来越多地与人类互动,理解人与AI之间的信任动态成为核心问题。尽管已有大量研究关注人类如何信任AI,但关于LLM代理如何有效信任人类的研究仍不足。在涉及贷款审批等场景中,LLM代理可能依赖某种隐式信任机制辅助决策。本文基于既有行为理论,提出方法研究大模型对人类可信度的判断,涵盖能力、善意和诚信三个维度,并考察人口统计变量的影响。通过针对五种主流语言模型在五种不同场景下的43,200次模拟实验,发现大模型的信任发展整体上与人类相似:多数情况下,信任强烈依赖于可信度,但在某些情境下也受年龄、宗教和性别影响,尤其在金融类场景中更为明显。该现象在文献常见任务和较新模型中尤为突出。尽管总体模式符合人类信任形成机制,但不同模型在信任评估上存在差异,部分情况下可信度与人口因素对信任预测力较弱。这些发现呼吁加强对人工智能对人类信任机制的理解与监控,以防止在敏感应用中出现非预期且有害的结果。
原文摘要 · Abstract (English)
As large language models (LLMs) and LLM-based agents increasingly interact with humans in decision-making contexts, understanding the trust dynamics between humans and AI agents becomes a central concern. While considerable literature studies how humans trust AI agents, it is much less understood how LLM-based agents develop effective trust in humans. LLM-based agents likely rely on some sort of implicit effective trust in trust-related contexts (e.g., evaluating individual loan applications) to assist and affect decision making. Using established behavioral theories, we develop an approach that studies whether LLMs trust depends on the three major trustworthiness dimensions: competence, benevolence and integrity of the human subject. We also study how demographic variables affect effective trust. Across 43,200 simulated experiments, for five popular language models, across five different scenarios we find that LLM trust development shows an overall similarity to human trust development. We find that in most, but not all cases, LLM trust is strongly predicted by trustworthiness, and in some cases also biased by age, religion and gender, especially in financial scenarios. This is particularly true for scenarios common in the literature and for newer models. While the overall patterns align with human-like mechanisms of effective trust formation, different models exhibit variation in how they estimate trust; in some cases, trustworthiness and demographic factors are weak predictors of effective trust. These findings call for a better understanding of AI-to-human trust dynamics and monitoring of biases and trust development patterns to prevent unintended and potentially harmful outcomes in trust-sensitive applications of AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。