arXiv:2503.15128cs.CLcs.AI2025-03被引 11

提升多语言机器生成文本检测器的抗干扰和泛化能力

Increasing the Robustness of the Fine-tuned Multilingual Machine-Generated Text Detectors

  • 通过鲁棒微调增强检测模型对文本伪装的抵抗力
  • 在跨语言和分布外数据上显著提升检测准确率
  • 适合需要可靠内容审核的平台与研究者使用

随着大语言模型(LLMs)的普及,其被用于生成有害内容的风险引发关注。近期研究证实了大语言模型存在漏洞,且极易被滥用。如今,人类已难以区分高质量机器生成文本与真实人工文本。因此,开发自动化检测机制至关重要,有助于识别网络信息中的机器生成内容,从而评估其可信度。本文提出一种针对检测任务的鲁棒微调方法,使检测模型更抗文本混淆,并在分布外数据上更具泛化能力。

原文摘要 · Abstract (English)

Since the proliferation of LLMs, there have been concerns about their misuse for harmful content creation and spreading. Recent studies justify such fears, providing evidence of LLM vulnerabilities and high potential of their misuse. Humans are no longer able to distinguish between high-quality machine-generated and authentic human-written texts. Therefore, it is crucial to develop automated means to accurately detect machine-generated content. It would enable to identify such content in online information space, thus providing an additional information about its credibility. This work addresses the problem by proposing a robust fine-tuning process of LLMs for the detection task, making the detectors more robust against obfuscation and more generalizable to out-of-distribution data.

文本检测大模型安全多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。