不依赖模型,用分位数曲线评估仿真器与真实世界的差距。
Model-Free Assessment of Simulator Fidelity via Quantile Curves
- 通过构建潜在参数置信集,间接估计仿真与真实系统的差异。
- 在世界价值观基准数据集上验证,四款大模型表现差异明显。
- 适用于分类、连续多维等复杂输出,适合做风险评估和模型对比。
随着生成式AI越来越多用于模拟现实系统,量化‘仿真到现实’的差距至关重要。对于每个感兴趣的输入场景(如调查问题或运行条件),真实系统与仿真系统对应未观测的潜在总体参数,其差异随场景变化。核心挑战在于:对任意场景,该差异无法直接观测,因为两个系统仅可通过有限样本获取,且样本量常异质分布。标准预测推断方法因此不适用,因其衡量的是可观测输出的不确定性,而非潜在总体参数的不确定性。为此,我们构建这些潜在参数的置信集,并以此推导出仿真与真实差异的稳健代理指标。进而估计该代理的分位数函数,获得仿真器在分布层面的风险图谱,支持多种统计摘要,包括新场景下真实输出分布的统计推断、条件风险价值(CVaR)等风险度量计算,以及不同仿真器间的严谨比较。本方法为模型无关,可处理一般输出空间,如分类调查响应和连续多维数据。我们通过评估四款主流大语言模型在WorldValueBench数据集上与人类群体的对齐程度,展示了该方法的实用价值。
原文摘要 · Abstract (English)
As generative AI models are increasingly used to simulate real-world systems, quantifying the ``sim-to-real'' gap is critical. For each input setting of interest -- which we call a \emph{scenario}, such as a survey question or operating condition -- the real and simulated systems are associated with unobserved latent population parameters, and their discrepancy varies across scenarios. A fundamental challenge is that, for any given scenario, this discrepancy cannot be observed directly, since both systems are accessible only through finite samples, often of heterogeneous sizes across scenarios. Standard predictive inference methods are therefore ill-suited, as they quantify uncertainty in observable outputs rather than latent population parameters. To address this, we construct confidence sets for these latent parameters and use them to derive a robust proxy for the sim-to-real discrepancy. We then estimate the quantile function of this proxy to obtain a distribution-level risk profile of the simulator, which supports a broad range of statistical summaries, including statistical inference for the real output distribution in a new scenario, the calculation of risk measures like Conditional Value-at-Risk (CVaR), and principled comparisons across simulators. Our method is model-agnostic and handles general output spaces, such as categorical survey responses and continuous multi-dimensional data. We demonstrate the practical utility of this method by evaluating the alignment of four major LLMs with human populations on the WorldValueBench dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。