微调大模型可生成难以分辨的人类级极化评论
Passing the Turing Test in Political Discourse: Fine-Tuning LLMs to Mimic Polarized Social Media Comments
- 用Reddit政治讨论数据微调开源大模型,生成立场一致的回应
- 模型输出在语言风格和情绪上与真人评论高度相似
- 适合关注AI伦理、虚假信息治理的研究者阅读
大型语言模型(LLMs)日益复杂,引发了其在在线环境中自动化生成具有说服力和偏见内容、加剧意识形态极化的担忧。本研究探讨了微调后的LLM在模拟和放大网络极化话语方面的程度。利用从Reddit提取的政治敏感讨论数据集,我们对一个开源大模型进行微调,使其生成上下文感知且立场一致的回复。通过语言学分析、情感评分和人工标注评估模型输出,重点关注其可信度和与原始话语的修辞一致性。结果表明,当在党派性数据上训练时,LLMs能够生成高度可信且具有挑衅性的评论,常与人类撰写的内容无法区分。这些发现引发了关于人工智能在政治言论、虚假信息及操纵活动中使用的重大伦理问题。论文最后讨论了对AI治理、平台监管以及检测对抗性微调风险工具开发的广泛影响。
原文摘要 · Abstract (English)
The increasing sophistication of large language models (LLMs) has sparked growing concerns regarding their potential role in exacerbating ideological polarization through the automated generation of persuasive and biased content. This study explores the extent to which fine-tuned LLMs can replicate and amplify polarizing discourse within online environments. Using a curated dataset of politically charged discussions extracted from Reddit, we fine-tune an open-source LLM to produce context-aware and ideologically aligned responses. The model's outputs are evaluated through linguistic analysis, sentiment scoring, and human annotation, with particular attention to credibility and rhetorical alignment with the original discourse. The results indicate that, when trained on partisan data, LLMs are capable of producing highly plausible and provocative comments, often indistinguishable from those written by humans. These findings raise significant ethical questions about the use of AI in political discourse, disinformation, and manipulation campaigns. The paper concludes with a discussion of the broader implications for AI governance, platform regulation, and the development of detection tools to mitigate adversarial fine-tuning risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。