通过多层级上下文关系建模,提升机器生成文本的检测准确率。
Multi-Level Contextual Token Relation Modeling for Machine-Generated Text Detection

- 设计轻量级马尔可夫校准模块,修正局部词元偏差。
- 引入规则支持推理模块,利用上下文统计规律增强全局判断。
- 在跨模型、跨领域场景下表现优异,计算开销低。
机器生成文本(MGT)带来虚假信息和网络钓鱼等风险,亟需可靠检测方法。基于度量的方法通过提取统计上可区分的特征,比易过拟合的复杂模型更实用。我们首次将代表性度量方法置于统一框架中,清晰评估其优劣。分析发现核心挑战:词元级检测分数易受生成过程固有随机性干扰。理论上推导了词元级分数的多跳转移特性,并探索其局部与全局关系。据此提出多层级上下文词元关系建模框架:局部关系通过轻量级马尔可夫校准模块,在聚合前修正词元证据;全局关系通过规则支持推理模块,利用上下文分数统计导出的显式逻辑规则。最终在联合多层级推理框架中融合局部校准分数与全局规则支持信号。大量实验表明,该方法在多种真实场景(包括跨大模型、跨领域)中均实现显著且广泛的性能提升,且计算开销低。
原文摘要 · Abstract (English)
Machine-generated texts (MGTs) pose risks such as disinformation and phishing, underscoring the need for reliable detection. Metric-based methods, which extract statistically distinguishable features of MGTs, are often more practical than complex model-based methods that are prone to overfitting. Given their diverse designs, we first place representative metric-based methods within a unified framework, enabling a clear assessment of their advantages and limitations. Our analysis identifies a core challenge across these methods: the token-level detection score is easily biased by the inherent randomness of the MGTs generation process. Then, we theoretically derive the multi-hop transitions of the token-level detection score and explore their local and global relations. Based on these findings, we propose a multi-level contextual token relation modeling framework for MGT detection. Specifically, for local relations, we model them through a lightweight Markov-informed calibration module that refines token-level evidence before aggregation. For global relations, we introduce a rule-support reasoning module that uses explicit logical rules derived from contextual score statistics. Finally, we combine the local calibrated score and the global rule-support reasoning signal in a joint multi-level inference framework. Extensive experiments show broad and substantial improvements across various real-world scenarios, including cross-LLM and cross-domain settings, with low computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。