arXiv:2512.08939cs.HCcs.AI2025-12被引 1

用大模型模拟医疗信任,发现其反应过于集中,细节差异捕捉不足。

Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust

  • 基于真实人群数据评估大模型数字孪生的医疗信任模拟表现。
  • 模拟结果方差小,极端选项选择少,整体分布更集中(所有p<0.001)。
  • 对教育程度等细微差异敏感度低,适合宏观趋势模拟而非精细分析。

大型语言模型(LLM)驱动的人类数字孪生作为新兴工具,在医疗系统研究中展现出巨大潜力。然而,其在模拟复杂人类心理特征(如对医疗系统的不信任)方面的实际能力尚不明确,这直接影响医护人员对基于LLM的人工智能系统在日常工作中辅助决策的信任与使用。本研究基于Twin-2K-500数据集,利用健康医疗系统不信任量表(HCSDS)对由人机数字孪生生成的模拟响应进行系统评估,采用已建立的人类受试样本,分析项目级分布、汇总统计及人口学亚组模式。结果显示,数字孪生生成的响应显著更集中,方差更低,且极端选项选择更少(所有p<0.001)。尽管数字孪生在年龄、性别等主要人口学模式上能大致复现人类行为,但在教育水平等次要差异上表现出较低敏感性。研究表明,当前基于LLM的数字孪生具备模拟群体趋势的潜力,但难以在个体细分群体中做出精准区分。该研究提示,现有数字孪生在建模复杂人类态度方面存在局限,需经过仔细校准与验证后方可用于健康系统工程中的推断分析或政策模拟。未来研究有必要考察大模型的情感推理机制,尤其在涉及社会敏感话题(如人机信任)的模拟中。

原文摘要 · Abstract (English)

Serving as an emerging and powerful tool, Large Language Model (LLM)-driven Human Digital Twins are showing great potential in healthcare system research. However, its actual simulation ability for complex human psychological traits, such as distrust in the healthcare system, remains unclear. This research gap particularly impacts health professionals' trust and usage of LLM-based Artificial Intelligence (AI) systems in assisting their routine work. In this study, based on the Twin-2K-500 dataset, we systematically evaluated the simulation results of the LLM-driven human digital twin using the Health Care System Distrust Scale (HCSDS) with an established human-subject sample, analyzing item-level distributions, summary statistics, and demographic subgroup patterns. Results showed that the simulated responses by the digital twin were significantly more centralized with lower variance and had fewer selections of extreme options (all p<0.001). While the digital twin broadly reproduces human results in major demographic patterns, such as age and gender, it exhibits relatively low sensitivity in capturing minor differences in education levels. The LLM-based digital twin simulation has the potential to simulate population trends, but it also presents challenges in making detailed, specific distinctions in subgroups of human beings. This study suggests that the current LLM-driven Digital Twins have limitations in modeling complex human attitudes, which require careful calibration and validation before applying them in inferential analyses or policy simulations in health systems engineering. Future studies are necessary to examine the emotional reasoning mechanism of LLMs before their use, particularly for studies that involve simulations sensitive to social topics, such as human-automation trust.

数字孪生大模型医疗信任心理模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。