用多种模型区分真假新闻,发现AI生成内容风格更统一。
Human vs. Machine Deception: Distinguishing AI-Generated and Human-Written Fake News Using Ensemble Learning
- 结合句法、词汇多样性等特征构建文本表示
- 可准确识别95%以上的人类与AI生成假新闻
- 适合关注信息真实性与模型检测的研究者
大型语言模型的广泛应用催生了新型人工智能生成的虚假新闻,与传统人工编造的谣言共存,引发对其差异及可辨别性的关注。本研究分析人类撰写与AI生成假新闻在语言、结构和情感上的差异,并评估机器学习与集成方法的区分能力。通过句法结构、词汇多样性、标点模式、可读性指数及恐惧、愤怒、喜悦、悲伤、信任、预期等情绪特征构建文档级特征表示。采用逻辑回归、随机森林、支持向量机、极端梯度提升和神经网络等多类分类模型,结合集成框架融合各模型预测结果。以准确率和受试者工作特征曲线下面积(AUC)评估性能。结果显示,模型表现优异且稳定,可读性特征为最具判别力的指标,AI生成文本呈现更一致的风格模式。集成学习相较于单一模型带来小幅但稳定的提升。研究表明,文本的风格与结构特性可有效区分人工智能生成的虚假信息与人工制造的假新闻。
原文摘要 · Abstract (English)
The rapid adoption of large language models has introduced a new class of AI-generated fake news that coexists with traditional human-written misinformation, raising important questions about how these two forms of deceptive content differ and how reliably they can be distinguished. This study examines linguistic, structural, and emotional differences between human-written and AI-generated fake news and evaluates machine learning and ensemble-based methods for distinguishing these content types. A document-level feature representation is constructed using sentence structure, lexical diversity, punctuation patterns, readability indices, and emotion-based features capturing affective dimensions such as fear, anger, joy, sadness, trust, and anticipation. Multiple classification models, including logistic regression, random forest, support vector machines, extreme gradient boosting, and a neural network, are applied alongside an ensemble framework that aggregates predictions across models. Model performance is assessed using accuracy and area under the receiver operating characteristic curve. The results show strong and consistent classification performance, with readability-based features emerging as the most informative predictors and AI-generated text exhibiting more uniform stylistic patterns. Ensemble learning provides modest but consistent improvements over individual models. These findings indicate that stylistic and structural properties of text provide a robust basis for distinguishing AI-generated misinformation from human-written fake news.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。