arXiv:2609.07305cs.CLcs.HC2026-09综述

用真实数据验证合成问卷面板,发现匹配率高不等于模拟出真实用户。

Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure

论文配图:Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure
图 1 · 摘自论文原文
  • 通过响应契约分析发现,合成群体选项空缺率达66/128,远高于真实人群。
  • 在八组对比中,概率模型使边际误差降低4.53至7.30点。
  • 匹配度无法区分真实用户模拟与直接人口估计,验证方式有缺陷。

人口统计学合成问卷面板常通过与公开调查的汇总结果匹配来验证。本文在来自三个国家四个机构的六组多选题中检验这一标准的有效性。核心分析聚焦于三组与真实目标群体共享人口框架的题目;其余为敏感性分析。响应契约主导了测量拟合度:在对齐的题目中,承诺型样本使128个选项槽中有66个为空(最多500人),而每选项独立概率抽取则全为0。在八组无上限模型对比中,概率方法使选项边际平均绝对误差(MAE)降低4.53至7.30点。有上限的题目在两个模型上出现反向,需投影至声明最大值才逆转。这些是测量效应:真实受访者为全选实现实例,而向量为潜在包含概率。公布的边际一致性也无法区分用户模拟与直接人口估计。在九组对齐组合中,无人物画像的人口比例查询平均误差为6.27,优于承诺型面板的12.39,且全部胜出。约束感知概率向量平均误差5.34,在四组中击败查询,表明基准挑战了现有验证标准,而非证明直接估计始终最优。在三组未公开的人口子集上,两种方法均不如直接复述全国分布。因此,人口边际一致性仅反映抽样契约和可得估计量,而非个体模拟证据。

原文摘要 · Abstract (English)

Demographic synthetic survey panels are often validated by matching aggregate answers to published surveys. We test what that certificate establishes across six multiselect batteries from four survey organisations in three countries. The headline analysis is restricted to three instruments whose synthetic cohort and human target share the stated population frame; three other batteries remain sensitivity analyses. The response contract dominates measured fidelity. In the aligned instruments, committed sets leave 66 of 128 model-battery option slots empty in panels of up to 500 respondents, versus 0 of 128 under per-option probability elicitation. Across eight uncapped model-instrument comparisons, probabilities reduce option-marginal MAE by 4.53 to 7.30 points. The capped instrument reverses on two models until the vectors are projected onto its stated maximum. These are measurement effects: human targets are realised check-all responses, whereas the vectors are latent inclusion propensities. Published marginal agreement also fails to discriminate respondent simulation from direct population estimation. On nine aligned model-battery pairs, a no-persona population-prevalence query averages 6.27 MAE versus 12.39 for committed panels and wins all nine comparisons. Constraint-aware probability vectors average 5.34 and beat the query on four of nine, so the baseline challenges the validation criterion rather than proving direct estimation uniformly best. On three unpublished demographic cells, neither approach beats reciting the national distribution. Population-marginal agreement is therefore evidence about an elicitation contract and an estimand obtainable without simulated respondents, not evidence of individual simulation.

合成数据问卷验证模拟用户边际拟合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。