通过可控内容重叠实验,揭示风格分类器依赖内容线索的机制。
Style or Content? Evaluating Style Classifiers with Controlled Content Overlap

- 设计参数α控制不同风格间的内容共享程度,量化内容与风格的相关性。
- 低重叠模型移除内容线索后性能下降,高重叠模型则更鲁棒。
- 内容可恢复性随α增加而降低,表明模型逐步学习到风格特征。
风格分类器可能利用与风格标签相关的自然数据内容线索,但我们缺乏系统评估其依赖程度的方法。本文基于平行圣经翻译构建受控内容重叠实验,定义重叠参数α为内容身份与风格标签间互信息的归一化残差,衡量风格类别间共享内容的比例:从无共享(α=0)到完全共享(α=1)。对基于RoBERTa的分类器进行跨重叠评估发现,低重叠模型在移除内容线索后性能下降明显,而高重叠模型具有更强的迁移鲁棒性。进一步的跨风格内容检索探针显示,随着α增大,内容可恢复性降低,训练动态也表明该去除过程是渐进的。结果表明,受控重叠提供了一种简单有效的诊断工具,用于区分风格学习与内容捷径。
原文摘要 · Abstract (English)
Style classifiers can use content cues that correlate with style labels in naturally collected data, yet we lack a systematic way to measure this reliance. We study this problem with a controlled content overlap setup built on parallel Bible translations. Specifically, we define the overlap parameter $α$ as the normalized residual of mutual information between content identity and style label, so that it measures how much content is shared across style classes: from no shared content ($α=0$) to fully shared content ($α=1$). Cross-overlap evaluation of RoBERTa-based classifiers shows that low-overlap models degrade when content cues are removed, while high-overlap models transfer more robustly. A cross-style content retrieval probe further shows that content becomes less recoverable as $α$ increases, with training dynamics showing this removal occurs gradually. Together, these results suggest that controlled overlap provides a simple diagnostic for separating style learning from content shortcuts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。