构建首个支持精确定位AI写作段落的多语言语料库
LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text
- 用现代大模型生成中英双语文本,覆盖混合作者场景
- 提供字符级标注,实现对AI生成内容的精准定位
- 助力开发更精细的AI文本检测模型,适合研究者使用
大型语言模型生成的类人文本日益普及,亟需可靠的检测系统。但当前进展受限于高质量训练数据的缺乏:现有数据集多基于过时模型、以英语为主,且未涵盖常见的混合人类与AI写作场景。尤其关键的是,尽管部分数据集涉及混合写作,却均无字符级标注,无法实现对AI生成片段的精确识别。为此,我们提出LLMTrace,一个大规模、双语(英语与俄语)的AI生成文本语料库。该数据集采用多种现代专有及开源大模型构建,旨在支持两大任务:传统全篇二分类(人类/人工智能)和借助字符级标注实现的AI生成区间检测。我们相信LLMTrace将为下一代更细致、实用的AI检测模型训练与评估提供关键资源。项目页面见 https://sweetdream779.github.io/LLMTrace-info/。
原文摘要 · Abstract (English)
The widespread use of human-like text from Large Language Models (LLMs) necessitates the development of robust detection systems. However, progress is limited by a critical lack of suitable training data; existing datasets are often generated with outdated models, are predominantly in English, and fail to address the increasingly common scenario of mixed human-AI authorship. Crucially, while some datasets address mixed authorship, none provide the character-level annotations required for the precise localization of AI-generated segments within a text. To address these gaps, we introduce LLMTrace, a new large-scale, bilingual (English and Russian) corpus for AI-generated text detection. Constructed using a diverse range of modern proprietary and open-source LLMs, our dataset is designed to support two key tasks: traditional full-text binary classification (human vs. AI) and AI-generated interval detection, facilitated by character-level annotations. We believe LLMTrace will serve as a vital resource for training and evaluating the next generation of more nuanced and practical AI detection models. The project page is available at https://sweetdream779.github.io/LLMTrace-info/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。