arXiv:2511.01192cs.CL2025-11

提出DEER框架,提升机器生成文本检测的泛化能力。

DEER: Disentangled Mixture of Experts with Instance-Adaptive Routing for Generalizable Machine-Generated Text Detection

  • 将领域特异与共享知识分离,由强化学习路由动态选择专家路径。
  • 在域内和域外数据上分别提升1.28%和2.92%的F1值。
  • 适合需要跨域稳定检测的开放世界应用场景。

随着大模型快速发展,机器生成文本检测面临严峻挑战,现有检测器在域迁移下性能显著下降。通过系统性初步研究,我们发现当前泛化策略存在两大根本缺陷:多域训练中领域特异性知识保留不全,以及推理时知识检索与检测目标不匹配。为此,我们提出DEER——一种解耦混合专家框架,将领域局部与领域不变知识显式分离至专用专家模块。不同于静态域匹配,DEER采用强化学习驱动的路由器,根据实例级检测奖励动态选择专家路径。该任务对齐、域无关机制通过优先考虑检测效用而非风格相似性,实现对未见分布的稳健适应。大量实验表明,DEER持续优于现有最先进检测器,在域内与域外数据集上分别实现1.28%和2.92%的平均F1提升,准确率分别提高1.35%和2.26%,为开放世界部署提供可靠泛化能力。

原文摘要 · Abstract (English)

Detecting machine-generated text has become a critical challenge amid the rapid advancement of LLMs, yet existing detectors degrade severely under domain shift. Through systematic pilot studies, we trace this vulnerability to two fundamental flaws in current generalization strategies, namely the incomplete preservation of domain-specific knowledge during multi-domain training and the misalignment between knowledge retrieval and the detection objective at inference. To address these gaps, we propose DEER, a Disentangled mixturE-of-ExpeRts framework that explicitly decouples domain-local and domain-invariant knowledge into specialized expert modules. Instead of static domain matching, DEER employs a reinforcement learning-driven router that selects expert pathways based on instance-level detection rewards. This task-aligned, domain-agnostic mechanism ensures robust adaptation to unseen distributions by prioritizing detection utility over stylistic resemblance. Extensive experiments demonstrate that DEER consistently outperforms state-of-the-art detectors, achieving average F1 improvements of 1.28% and 2.92%, and accuracy gains of 1.35% and 2.26% on in-domain and out-of-domain datasets, offering reliable generalization for open-world deployment.

文本检测泛化能力专家混合强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。