arXiv:2411.16813cs.CLcs.AI2024-11

用政治辩论数据微调大模型,发现越文明的数据越僵化,跨平台训练反而更毒。

Incivility and Rigidity: Evaluating the Risks of Fine-Tuning LLMs for Political Argumentation

  • 用推特和Reddit两类政治言论数据微调GPT-3.5 Turbo,对比其生成论据的风格差异。
  • Reddit数据微调后论点安全但缺乏灵活性,跨平台训练加剧攻击性与毒性。
  • 提出四维度评价框架,适合开发文明对话辅助系统的研究者参考。

推特(现为X)和Reddit等平台上的不文明言论,给构建能支持理性政治辩论的AI系统带来挑战。本文对基于美国国会推文回复(高不文明)和Reddit的r/ChangeMyView版块(低不文明)的两组政治言论数据,对GPT-3.5 Turbo进行微调实验。评估聚焦于数据构成与提示策略如何影响模型生成论据的修辞框架与讨论质量。结果显示,经Reddit数据微调的模型生成论据更安全但修辞僵化;跨平台微调则加剧敌意语气与毒性。基于提示的引导可减少明显攻击性(如人身攻击),但无法完全抵消噪声训练数据的影响。本文提出一套涵盖论证理由、互惠性、一致性与权威性的修辞评估量表,并提供针对内容创作、审核与协商支持系统的实现指南。

原文摘要 · Abstract (English)

Incivility on platforms such as Twitter (now X) and Reddit complicates the development of AI systems that can support productive, rhetorically sound political argumentation. We present experiments with \textit{GPT-3.5 Turbo} fine-tuned on two contrasting datasets of political discourse: high-incivility Twitter replies to U.S. Congress and low-incivility posts from Reddit's \textit{r/ChangeMyView}. Our evaluation examines how data composition and prompting strategies affect the rhetorical framing and deliberative quality of model-generated arguments. Results show that Reddit-finetuned models generate safer but rhetorically rigid arguments, while cross-platform fine-tuning amplifies adversarial tone and toxicity. Prompt-based steering reduces overt toxicity (e.g., personal attacks) but cannot fully offset the influence of noisy training data. We introduce a rhetorical evaluation rubric - covering justification, reciprocity, alignment, and authority - and provide implementation guidelines for authoring, moderation, and deliberation-support systems.

大模型政治辩论微调风险修辞评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。