开源工具VietAIDetector可零样本检测越南语AI文本,支持长文本和扫描文档。
VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text

- 基于越南语专用模型的零样本检测,无需领域训练数据。
- 在跨域数据集上表现优于主流英文检测方法。
- 提供网页界面与灵活阈值选择,适合内容审核与学术验证。
近年来,区分AI生成文本与人工写作仍具挑战。本文提出VietAIDetector,一款专为检测越南语AI生成文本设计的开源工具。用户可通过Gradio网页界面输入原始越南语文本或常见文件格式,包括扫描文档和超出大语言模型上下文长度限制的超长文本。该工具核心采用零样本检测方法,无需领域特定训练数据,建立在先前VietBinoculars与Binoculars研究基础上。其依托越南语专用语言模型,在跨域数据集上评估,性能优于主要针对英语开发的现有方法。用户可根据F1分数、准确率或[email protected]需求选择最优检测阈值。结果通过网页界面呈现,支持可疑文本的快速审查与PDF报告下载。工具已公开发布于https://github.com/trieuntu/VietAIDetector。
原文摘要 · Abstract (English)
In recent years, distinguishing between AI-generated text and human-written text has remained a challenge. In this paper, we introduce VietAIDetector, an open-source tool designed specifically for detecting Vietnamese AI-generated text. It allows users to interact through a Gradio web interface with inputs ranging from raw Vietnamese text to common text file formats, including scanned documents and exceptionally long texts that exceed the context size of the employed Large Language Models (LLMs). The core component of the tool employs a Zero-Shot approach to detect AI-generated text without requiring domain-specific training data, building upon the previous VietBinoculars and Binoculars research. The tool is built upon a Vietnamese-specific language model and has been evaluated on out-of-domain datasets, demonstrating superior performance compared to existing methods primarily developed for English. Additionally, users can select optimal detection thresholds based on F1 score, accuracy, or [email protected] requirements. The results are presented through the web interface, allowing users to easily review and verify suspicious texts or download them as a PDF report. The tool is publicly available at https://github.com/trieuntu/VietAIDetector
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。