arXiv:2503.09289cs.CLcs.AI2025-03中稿 · NAACL综述被引 5

检测泰米尔语和马拉雅拉姆语的AI生成评论,提升低资源语言平台可信度

Unmask It! AI-Generated Product Review Detection in Dravidian Languages

  • 对比多种机器学习与Transformer模型,评估其在低资源语言中的检测能力
  • 基于Indic-BERT等先进模型,对AI生成评论识别准确率显著提升
  • 为南印度语言电商评论安全提供可落地的检测方案,适合平台风控使用

生成式AI的兴起导致大量AI生成评论涌现,严重威胁在线平台的可信度。评论是消费者决策的关键信息来源,虚假内容会误导用户、破坏信任并助长数字市场欺诈。本研究聚焦泰米尔语和马拉雅拉姆语这两种低资源语言的AI生成评论检测,比较了传统机器学习方法与先进Transformer模型(如Indic-BERT、IndicSBERT、MuRIL、XLM-RoBERTa和MalayalamBERT)的性能。结果表明,利用前沿Transformer模型能有效识别AI生成内容,在低资源语言场景下展现出显著检测潜力。

原文摘要 · Abstract (English)

The rise of Generative AI has led to a surge in AI-generated reviews, often posing a serious threat to the credibility of online platforms. Reviews serve as the primary source of information about products and services. Authentic reviews play a vital role in consumer decision-making. The presence of fabricated content misleads consumers, undermines trust and facilitates potential fraud in digital marketplaces. This study focuses on detecting AI-generated product reviews in Tamil and Malayalam, two low-resource languages where research in this domain is relatively under-explored. We worked on a range of approaches - from traditional machine learning methods to advanced transformer-based models such as Indic-BERT, IndicSBERT, MuRIL, XLM-RoBERTa and MalayalamBERT. Our findings highlight the effectiveness of leveraging the state-of-the-art transformers in accurately identifying AI-generated content, demonstrating the potential in enhancing the detection of fake reviews in low-resource language settings.

AI检测低资源语言评论伪造NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。