arXiv:2609.06646cs.CVcs.AI2026-09

笑声起始时间标注存在系统性差异,影响评测可靠性。

When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization

论文配图:When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization
图 1 · 摘自论文原文
  • 发现标注者对笑声起止点的分歧具有规律性,非随机噪声。
  • 笑声起始点分歧是结束点的1.73倍,短笑声分歧更普遍(77%)。
  • 提出基于标注分布的校准评估方法,提升评测公平性。

标注者在笑声边界和细微笑声上普遍存在分歧,但这种分歧并非随机噪声,而是具有系统性模式。我们在SMILE-Temporal基准(672个视频,1,683个事件)上重新标注,每个视频由3-5名标注者完成(一致性系数α=0.757),发现:结束点的分歧是起始点的1.73倍;短笑声的分歧率(77%)远高于完整笑声(20%);且分歧可由事件属性预测(AUC=0.831)。以单一标注为标准的评测存在偏差:系统F1分数因参考标注不同波动0.246,正确排序系统仅69.7%(对比所有标注时为80%)。为此提出校准评估方法,使用符合校准的容忍带(结束点0.727秒,起始点0.5秒),对齐标注分布。标注数据与分析代码已公开于https://github.com/WSCSports/SMILE-Disagreement。

原文摘要 · Abstract (English)

Annotators routinely disagree on laughter boundaries and subtle chuckles, yet temporal laughter localization typically evaluates against a single reference annotation. We show that this disagreement is structured rather than random noise. Re-annotating the SMILE-Temporal benchmark (672 videos, 1,683 events) with 3-5 annotators per video (alpha = 0.757), we find systematic patterns: disagreement is 1.73 times larger at offsets than onsets, far more common for chuckles than full laughs (77% vs. 20%), and predictable from event attributes (AUC = 0.831). Evaluating against a single annotator breaks down under this structure: system scores shift by 0.246 F1 depending on the chosen ground truth, correctly ranking systems only 69.7% of the time (vs. 80% against all annotators). We propose a disagreement-calibrated evaluation that scores predictions against the full annotator distribution using conformally calibrated tolerance bands (wider at offsets, 0.727s, than onsets, 0.5s). The per-annotator annotations and analysis code are available at https://github.com/WSCSports/SMILE-Disagreement .

音频分析标注差异评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。