arXiv:2601.10719cs.AIcs.CL2026-01

发现大模型隐式学习了人类信任的心理机制,无需刻意训练即可识别可信内容。

Do You Trust Me? Cognitive-Affective Signatures of Trustworthiness in Large Language Models

  • 通过认知评估与情绪标注数据,分析模型在不同层和注意力头的激活差异
  • 高可信与低可信文本的激活模式系统性不同,且公平性等维度关联最强
  • 结果可直接用于提升AI系统的可信度,适合研究人机信任与AI透明性的人看

用户对在线信息的信任感决定了其信息行为,但当前大型语言模型(如 Llama 3.1 8B、Qwen 2.5 7B、Mistral 7B)是否以心理一致的方式表征这一概念仍不明确。本研究基于标注了认知评估、情绪与行为意图的 PEACE-Reviews 数据集,分析指令微调后的模型如何在网页式叙述中编码感知可信度。结果显示,不同层级与注意力头的激活差异能系统区分高可信与低可信文本,表明信任线索在预训练阶段即被隐式编码。探测分析显示,信任信号可线性解码,微调仅优化而非重构这些表征。最强关联出现在公平性、确定性与自我问责等人类信任核心维度。这表明现代 LLM 在无显式监督下内化了心理上合理的信任信号,为构建可信、透明、可靠的网络生态中人工智能系统提供了表征基础。代码与附录见:https://github.com/GerardYeo/TrustworthinessLLM。

原文摘要 · Abstract (English)

Perceived trustworthiness underpins how users navigate online information, yet it remains unclear whether large language models (LLMs),increasingly embedded in search, recommendation, and conversational systems, represent this construct in psychologically coherent ways. We analyze how instruction-tuned LLMs (Llama 3.1 8B, Qwen 2.5 7B, Mistral 7B) encode perceived trustworthiness in web-like narratives using the PEACE-Reviews dataset annotated for cognitive appraisals, emotions, and behavioral intentions. Across models, systematic layer- and head-level activation differences distinguish high- from low-trust texts, revealing that trust cues are implicitly encoded during pretraining. Probing analyses show linearly de-codable trust signals and fine-tuning effects that refine rather than restructure these representations. Strongest associations emerge with appraisals of fairness, certainty, and accountability-self -- dimensions central to human trust formation online. These findings demonstrate that modern LLMs internalize psychologically grounded trust signals without explicit supervision, offering a representational foundation for designing credible, transparent, and trust-worthy AI systems in the web ecosystem. Code and appendix are available at: https://github.com/GerardYeo/TrustworthinessLLM.

大模型信任建模认知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。