arXiv:2411.04950cs.CL2024-11被引 1

提出新方法检测文本主题连续性对分类结果的干扰,提升判别可靠性。

Estimating the Influence of Sequentially Correlated Literary Properties in Textual Classification: A Data-Centric Hypothesis-Testing Approach

  • 构建基于序列依赖的替代标签生成框架,保留上下文关联
  • 发现监督模型易将主题相似误判为风格差异,导致假阳性
  • 适用于作者鉴定与敏感文本分析,尤其能减少误判

我们提出一种数据驱动的假设检验框架,用于量化文本中主题连贯性等序列相关文学特征对分类任务的影响。通过将标签序列建模为随机过程,并利用经验自协方差矩阵生成保持序列依赖的替代标签,实现统计检验,判断分类结果是受主题结构驱动,还是由非序列特征(如写作风格)主导。在英语散文语料库上,对比了传统(词n-gram、字符k-mer)和神经(对比训练)嵌入在有监督与无监督分类中的表现。关键发现:监督及神经模型更易受序列相关性干扰,产生假阳性(将主题共性误认为风格信号)。而使用传统特征的无监督模型在同题材场景下常呈现高真阳性率且假阳性极低。该方法可有效分离序列与非序列影响,为作者归属、法医语言学及匿名/合成文本分析提供可信评估依据,强调控制序列相关性对降低误判、确保分类结果反映真实风格差异至关重要。

原文摘要 · Abstract (English)

We introduce a data-centric hypothesis-testing framework to quantify the influence of sequentially correlated literary properties--such as thematic continuity--on textual classification tasks. Our method models label sequences as stochastic processes and uses an empirical autocovariance matrix to generate surrogate labelings that preserve sequential dependencies. This enables statistical testing to determine whether classification outcomes are primarily driven by thematic structure or by non-sequential features like authorial style. Applying this framework across a diverse corpus of English prose, we compare traditional (word n-grams and character k-mers) and neural (contrastively trained) embeddings in both supervised and unsupervised classification settings. Crucially, our method identifies when classifications are confounded by sequentially correlated similarity, revealing that supervised and neural models are more prone to false positives--mistaking shared themes and cross-genre differences for stylistic signals. In contrast, unsupervised models using traditional features often yield high true positive rates with minimal false positives, especially in genre-consistent settings. By disentangling sequential from non-sequential influences, our approach provides a principled way to assess and interpret classification reliability. This is particularly impactful for authorship attribution, forensic linguistics, and the analysis of redacted or composite texts, where conventional methods may conflate theme with style. Our results demonstrate that controlling for sequential correlation is essential for reducing false positives and ensuring that classification outcomes reflect genuine stylistic distinctions.

文本分类主题分析作者识别模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。