arXiv:2512.21709cs.CLcs.AI2025-12中稿 · publication in 202…被引 1

首次系统检测孟加拉语AI改写文本,发现微调后模型准确率达91%

Detecting AI-Generated Paraphrases in Bengali: A Comparative Study of Zero-Shot and Fine-Tuned Transformers

  • 用五种Transformer模型对比零样本与微调效果
  • 微调后XLM-RoBERTa等模型准确率和F1达91%
  • 为孟加拉语反AI伪造提供可复用检测框架

大型语言模型能生成接近人类写作的文本,引发虚假信息和内容操纵担忧。现有研究已覆盖多种语言,但孟加拉语仍缺乏探索。其丰富的词汇与复杂结构使区分人写与AI生成文本尤为困难。本研究评估五种基于Transformer的模型:XLMRoBERTa-Large、mDeBERTaV3-Base、BanglaBERT-Base、IndicBERT-Base 和 MultilingualBERT-Base。零样本测试显示所有模型表现接近随机(约50%准确率),凸显任务微调的必要性。微调后,XLM-RoBERTa、mDeBERTa和MultilingualBERT在准确率和F1-score上均达到约91%;而IndicBERT表现较差,表明其在此任务上微调效果有限。本工作推进了孟加拉语中AI生成文本的检测,为构建抗伪系统奠定基础。

原文摘要 · Abstract (English)

Large language models (LLMs) can produce text that closely resembles human writing. This capability raises concerns about misuse, including disinformation and content manipulation. Detecting AI-generated text is essential to maintain authenticity and prevent malicious applications. Existing research has addressed detection in multiple languages, but the Bengali language remains largely unexplored. Bengali's rich vocabulary and complex structure make distinguishing human-written and AI-generated text particularly challenging. This study investigates five transformer-based models: XLMRoBERTa-Large, mDeBERTaV3-Base, BanglaBERT-Base, IndicBERT-Base and MultilingualBERT-Base. Zero-shot evaluation shows that all models perform near chance levels (around 50% accuracy) and highlight the need for task-specific fine-tuning. Fine-tuning significantly improves performance, with XLM-RoBERTa, mDeBERTa and MultilingualBERT achieving around 91% on both accuracy and F1-score. IndicBERT demonstrates comparatively weaker performance, indicating limited effectiveness in fine-tuning for this task. This work advances AI-generated text detection in Bengali and establishes a foundation for building robust systems to counter AI-generated content.

文本检测孟加拉语Transformer微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。