42%的LLM对话分析结论可能因忽略自相关而虚假,需修正统计方法。
The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious
- 发现对话轮次间存在强自相关,现有分析方法普遍忽略此问题。
- 42%的显著性结果在修正后不再成立,不同指标类别差异巨大。
- 提出双阶段校正框架,适合做对话评估的科研人员使用。
轮次级指标广泛用于评估多轮人机对话中的安全、谄媚度与对话质量等属性。然而,对话中连续轮次并非统计独立,当前几乎所有评估流程均未对此进行统计推断修正。本文系统分析了202个多轮对话(11,639个轮次对,5位德语使用者,4个LLM平台)中66个轮次级指标的自相关结构,发现朴素的合并分析会导致显著性估计严重膨胀:42%在标准合并检验下显著的关联,在聚类稳健校正后失效。自相关导致的膨胀在不同类别间差异显著,非记忆类指标(如热循环、帧距、滚动窗口等)合计达33%,个别类别最高达100%;而无记忆类指标(嵌入速度、方向性、差分等)合计仅14%。本文提出结合Chelton(1983)有效自由度与对话级块重抽样的两阶段校正框架,并在预注册留出集上验证:聚类稳健指标复制率达57%,远高于仅用合并法的30%。提供设计原则、发表检查清单及开源代码。对近30篇顶会论文的调查显示,仅4篇提及时间依赖性,26篇完全未校正。
原文摘要 · Abstract (English)
Turn-level metrics are widely used to evaluate properties of multi-turn human-LLM conversations, from safety and sycophancy to dialogue quality. However, consecutive turns within a conversation are not statistically independent -- a fact that virtually all current evaluation pipelines fail to correct for in their statistical inference. We systematically characterize the autocorrelation structure of 66 turn-level metrics across 202 multi-turn conversations (11,639 turn pairs, 5 German-speaking users, 4 LLM platforms) and demonstrate that naive pooled analysis produces severely inflated significance estimates: 42% of associations that appear significant under standard pooled testing fail to survive cluster-robust correction. The inflation varies substantially across categories rather than scaling linearly with autocorrelation: three memoryless families (embedding velocity, directional, differential) aggregate to 14%, while the seven non-memoryless families (thermo-cycle, frame distance, lexical/structural, rolling windows, cumulative, interaction, timestamp) aggregate to 33%, with individual category rates ranging from 0% to 100% depending on per-family effect size. We present a two-stage correction framework combining Chelton (1983) effective degrees of freedom with conversation-level block bootstrap, and validate it on a pre-registered hold-out split where cluster-robust metrics replicate at 57% versus 30% for pooled-only metrics. We provide concrete design principles, a publication checklist, and open-source code for the correction pipeline. A survey of ~30 recent papers at major NLP and AI venues that compute turn-level statistics in LLM evaluations finds that only 4 address temporal dependence at all, and 26 do not correct for it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。