arXiv:2605.01311cs.LGecon.EM2026-05

解决大模型生成评估中的选择偏差问题,提升真实性能判断准确性

The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice

论文配图:The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice
图 1 · 摘自论文原文
  • 用观测日志、随机实验与仿真回放三源数据融合分析
  • 小规模随机实验+仿真可准确还原模型因果性能,日志仅用于降噪
  • 适合关注模型真实表现评估的研究者与产品团队

基于使用日志的离线语言模型评估在模型选择存在混淆时会产生偏差:用户行为既影响选用哪个模型,也影响对输出的评分,导致原始分数比较混杂了自选群体。小规模随机实验可通过强制替换模型选择来打破此偏差,但实际应用中稀少且成本高。本文提出三源设计:利用大规模混淆观测日志(OBS)获取规模,结合小规模随机实验(EXP)获得无偏评分,以及离线模拟器(SIM)在缓存上下文中重放候选模型。核心结果为一个识别定理,证明随机实验与模拟器共同足以恢复因果模型值;观测日志仅用于后续降低估计误差,而非使因果比较成立。六种估计器在受控半合成验证及两个真实任务缓存基准(摘要与编程)中评估。无一种方法在所有场景占优,相对表现取决于无偏实验监督量及目标奖励与观测结构的一致性。

原文摘要 · Abstract (English)

Offline evaluation of language models from usage logs is biased when model choice is confounded: the same user-side factors that influence which model is used can also influence how its output is judged, so raw comparisons of logged scores mix self-selected populations rather than estimating a common quantity of interest. A small randomized experiment can break this bias by overriding model choice, but in practice such experiments are scarce and costly. We study a three-source design that combines a large confounded observational log (OBS) for scale, a small randomized experiment (EXP) for unconfounded scoring, and an offline simulator (SIM) that replays candidate models on cached contexts. Our main result is an identification theorem showing that the randomized experiment and the simulator are together enough to recover causal model values; the observational log enters only afterward, to reduce estimation error rather than to make the causal comparison valid. Six estimator families are evaluated in a controlled semi-synthetic validation and in two real-task cached benchmarks for summarization and coding. No family dominates every regime; relative performance depends on the amount of unbiased EXP supervision and on how closely the target reward aligns with OBS-derived structure.

模型评估因果推断大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。