arXiv:2502.21028cs.HCcs.AI2025-02被引 11

提出首个量化人对大模型信任的6项指标,揭示信任受性格与使用经验影响。

Measuring and identifying factors of individuals' trust in Large Language Models

  • 构建心理量表TILLMI,区分情感亲近与认知依赖两类信任
  • 1000人调查验证量表信效度,因子分析确认双维度结构
  • 年轻人、外向者更易信任大模型,无使用经验者信任度更低

大语言模型(LLMs)可进行类人对话,但关于人机交互中信任形成的研究仍很稀缺。本文提出信任-大模型指数(TILLMI),基于McAllister的认知与情感信任框架,设计并验证了一套心理测量量表。通过1000名美国受访者样本,采用基于大模型模拟的验证协议,探索性因子分析识别出双因子结构,剔除冗余项后得到6项量表。在独立子样本上,验证性因子分析显示良好拟合度(CFI=.995, TLI=.991, RMSEA=.046, p_X²>.05)。收敛效度分析表明,对大模型的信任与开放性、外向性、认知灵活性正相关,与神经质负相关。据此将两因子解释为‘与大模型的情感亲近’和‘对大模型的依赖程度’。年轻男性比年长女性更倾向于亲近与依赖大模型;无直接使用经验者信任水平低于使用者。研究为评估大模型语义交互中的信任提供实证基础,有助于负责任的设计与人机协同。

原文摘要 · Abstract (English)

Large Language Models (LLMs) can engage in human-looking conversational exchanges. Although conversations can elicit trust between users and LLMs, scarce empirical research has examined trust formation in human-LLM contexts, beyond LLMs' trustworthiness or human trust in AI in general. Here, we introduce the Trust-In-LLMs Index (TILLMI) as a new framework to measure individuals' trust in LLMs, extending McAllister's cognitive and affective trust dimensions to LLM-human interactions. We developed TILLMI as a psychometric scale, prototyped with a novel protocol we called LLM-simulated validity. The LLM-based scale was then validated in a sample of 1,000 US respondents. Exploratory Factor Analysis identified a two-factor structure. Two items were then removed due to redundancy, yielding a final 6-item scale with a 2-factor structure. Confirmatory Factor Analysis on a separate subsample showed strong model fit ($CFI = .995$, $TLI = .991$, $RMSEA = .046$, $p_{X^2} > .05$). Convergent validity analysis revealed that trust in LLMs correlated positively with openness to experience, extraversion, and cognitive flexibility, but negatively with neuroticism. Based on these findings, we interpreted TILLMI's factors as "closeness with LLMs" (affective dimension) and "reliance on LLMs" (cognitive dimension). Younger males exhibited higher closeness with- and reliance on LLMs compared to older women. Individuals with no direct experience with LLMs exhibited lower levels of trust compared to LLMs' users. These findings offer a novel empirical foundation for measuring trust in AI-driven verbal communication, informing responsible design, and fostering balanced human-AI collaboration.

大模型信任心理量表人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。