arXiv:2508.13768cs.CL2025-08AAAI被引 4

通过频域分析提升机器文本检测的跨领域泛化能力

MGT-Prism: Enhancing Domain Generalization for Machine-Generated Text Detection via Spectral Alignment

  • 从频域视角设计低频滤波与动态谱对齐,捕捉域不变特征
  • 在11个数据集上平均提升0.90%准确率和0.92%F1值
  • 适合需要跨领域通用的文本生成检测场景

大型语言模型生成的文本越来越接近人类写作风格。当前的机器生成文本(MGT)检测器在同领域内表现良好,但在不同领域间泛化能力差,主要因数据源间的领域偏移。本文提出MGT-Prism,从频域角度提升检测器的领域泛化能力。通过分析文本表示的频域特性,我们发现不同领域的文本存在一致的频谱模式,而机器生成与人类写作在幅度上差异显著。基于此,设计了低频域滤波模块以去除受领域偏移影响的文档级特征,并引入动态谱对齐策略提取任务相关且域不变的特征。大量实验表明,MGT-Prism在三个领域泛化场景下的11个测试数据集上,平均准确率提升0.90%,F1分数提升0.92%,优于现有最先进方法。

原文摘要 · Abstract (English)

Large Language Models have shown growing ability to generate fluent and coherent texts that are highly similar to the writing style of humans. Current detectors for Machine-Generated Text (MGT) perform well when they are trained and tested in the same domain but generalize poorly to unseen domains, due to domain shift between data from different sources. In this work, we propose MGT-Prism, an MGT detection method from the perspective of the frequency domain for better domain generalization. Our key insight stems from analyzing text representations in the frequency domain, where we observe consistent spectral patterns across diverse domains, while significant discrepancies in magnitude emerge between MGT and human-written texts (HWTs). The observation initiates the design of a low frequency domain filtering module for filtering out the document-level features that are sensitive to domain shift, and a dynamic spectrum alignment strategy to extract the task-specific and domain-invariant features for improving the detector's performance in domain generalization. Extensive experiments demonstrate that MGT-Prism outperforms state-of-the-art baselines by an average of 0.90% in accuracy and 0.92% in F1 score on 11 test datasets across three domain-generalization scenarios.

文本检测域泛化频域分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。