arXiv:2609.07987cs.AIstat.AP2026-09

数字孪生能否替代人类测量?关键看它能否支持有效推断。

When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability

论文配图:When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability
图 1 · 摘自论文原文
  • 提出统计可替代性标准,评估数字孪生是否可减少人类数据采集
  • 发现孪生模型能复现平均效应,但难识别个体差异
  • 强调应以推断有效性而非行为相似度评价数字孪生

基于大语言模型的数字孪生有望减少重复的人类数据收集,但现有评估缺乏证据证明其能否在保持有效推断的前提下降低人类测量需求。为此,我们引入统计可替代性这一推断准则,评估孪生预测在特定估计量上减少人类测量的能力。我们构建了一个基于混合主体与预测驱动推断的框架,从四个维度评估:总体保真度、配对响应者层面信号、有限样本下人类标签恢复能力、跨群体稳定性。在两项涵盖行为实验、多个模型和不同响应者表征的实证研究中,我们发现数字孪生可再现平均人类效应,但几乎无法提供关于个体偏离均值的信息。更先进的模型和更丰富的响应者信息提升了部分性能维度,但未稳定转化为人类数据节省。人类校准可降低总体预测误差,但有限标注样本常无法带来稳定精度提升。重要的是,行为保真度既非统计可替代性的必要条件,也非充分条件。整体而言,表明应以支持有效科学推断的能力来评估人工智能生成证据,而非仅关注其复制人类结果的能力。因此,数字孪生在确认性使用中应以是否减少对人类数量的不确定性为评判标准,而非仅看其是否复现人类均值、分布或效应。

原文摘要 · Abstract (English)

LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.

数字孪生统计推断大模型应用人机数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。