arXiv:2501.09813cs.CL2025-01

Qwen团队在多语言伪文本检测中取得领先,性能超越35支队伍。

Qwen it detect machine-generated text?

  • 融合掩码与因果语言模型,提升多语言生成文本识别能力。
  • 子任务A的F1微平均得分0.8333,排名第一;宏平均得分0.8301,排名第二。
  • 适合关注生成内容检测、多语言NLP应用的研究者与工程师。

本文介绍了布加勒斯特大学-自然语言处理团队在2025年COLING生成式AI研讨会任务1:二分类多语言机器生成文本检测中的方法。我们探索了掩码语言模型与因果语言模型的结合应用。在子任务A中,我们的最佳模型在F1微平均(辅助评分)上获得0.8333,位列36支参赛队伍第一;在F1宏平均(主评分)上得分为0.8301,位列第二。结果表明,该方法在跨语言生成文本检测任务中具有显著优势。

原文摘要 · Abstract (English)

This paper describes the approach of the Unibuc - NLP team in tackling the Coling 2025 GenAI Workshop, Task 1: Binary Multilingual Machine-Generated Text Detection. We explored both masked language models and causal models. For Subtask A, our best model achieved first-place out of 36 teams when looking at F1 Micro (Auxiliary Score) of 0.8333, and second-place when looking at F1 Macro (Main Score) of 0.8301

文本检测多语言生成内容Qwen

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。