用政治辩论数据测试大模型谬误分类,发现情感音频反而降低准确率
Joint Effects of Argumentation Theory, Audio Modality and Data Enrichment on LLM-Based Fallacy Classification
- 基于论证理论设计两种思维链提示框架
- 加入情感音调信息后谬误分类准确率下降12%
- 适合研究人机推理偏差与多模态输入影响的学者
本研究探讨上下文与情感基调元数据对大语言模型(LLM)在谬误分类任务中推理能力与性能的影响,聚焦于美国总统辩论场景。基于真实辩论数据,采用Qwen-3(8B)模型对六类谬误进行分类,对比不同提示策略的效果。引入两种理论驱动的思维链框架:普拉格马-辩证法与论证周期表,并在三种输入设置下评估:仅文本、文本+上下文、文本+上下文+基于音频的情感基调元数据。结果表明,尽管理论提示能提升可解释性,但加入上下文尤其是情感基调元数据后,模型性能普遍下降;情感特征使模型更倾向将陈述标记为“诉诸情感”谬误,削弱逻辑判断力。整体而言,基础提示往往优于增强版本,暗示额外输入可能引发注意力分散,反而损害谬误识别效果。
原文摘要 · Abstract (English)
This study investigates how context and emotional tone metadata influence large language model (LLM) reasoning and performance in fallacy classification tasks, particularly within political debate settings. Using data from U.S. presidential debates, we classify six fallacy types through various prompting strategies applied to the Qwen-3 (8B) model. We introduce two theoretically grounded Chain-of-Thought frameworks: Pragma-Dialectics and the Periodic Table of Arguments, and evaluate their effectiveness against a baseline prompt under three input settings: text-only, text with context, and text with both context and audio-based emotional tone metadata. Results suggest that while theoretical prompting can improve interpretability and, in some cases, accuracy, the addition of context and especially emotional tone metadata often leads to lowered performance. Emotional tone metadata biases the model toward labeling statements as \textit{Appeal to Emotion}, worsening logical reasoning. Overall, basic prompts often outperformed enhanced ones, suggesting that attention dilution from added inputs may worsen rather than improve fallacy classification in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。