对比GPT-5.2等大模型在论点分类中的表现,揭示其优劣与共性缺陷。
A comprehensive study of LLM-based argument classification: from Llama through DeepSeek to GPT-5.2
- 测试多款主流大模型,结合思维链、重述提示等策略提升分类效果。
- 最佳模型在Args.me数据集达91.9%准确率,整体性能提升2%-8%。
- 发现模型对隐含批评、复杂结构理解差,且易受提示形式影响。
论点挖掘(AM)是自动识别和分类论点成分(如主张与前提)及其关系的跨学科研究领域。大语言模型(LLMs)的进展显著提升了论点分类性能。本研究全面评估了GPT-5.2、Llama 4和DeepSeek等前沿大模型在Args.me和UKP等公开数据集上的表现,采用思维链提示、提示重述、投票及置信度分类等先进策略。通过定量指标与定性错误分析,发现最优模型(GPT-5.2)在UKP数据集上准确率达78.0%,在Args.me上达91.9%。提示重述、多提示投票与置信度估计使模型准确率和F1值普遍提升2%至8%。但定性分析显示,各模型普遍存在对提示形式敏感、难以识别隐含批评、解析复杂论证结构困难及与特定主张对齐偏差等问题。本研究首次结合量化评测与定性分析,在多个数据集上系统评估多种大模型与提示策略。
原文摘要 · Abstract (English)
Argument mining (AM) is an interdisciplinary research field focused on the automatic identification and classification of argumentative components, such as claims and premises, and the relationships between them. Recent advances in large language models (LLMs) have significantly improved the performance of argument classification compared to traditional machine learning approaches. This study presents a comprehensive evaluation of several state-of-the-art LLMs, including GPT-5.2, Llama 4, and DeepSeek, on large publicly available argument classification corpora such as Args.me and UKP. The evaluation incorporates advanced prompting strategies, including Chain-of- Thought prompting, prompt rephrasing, voting, and certainty-based classification. Both quantitative performance metrics and qualitative error analysis are conducted to assess model behavior. The best-performing model in the study (GPT-5.2) achieves a classification accuracy of 78.0% (UKP) and 91.9% (Args.me). The use of prompt rephrasing, multi-prompt voting, and certainty estimation further improves classification performance and robustness. These techniques increase the accuracy and F1 metric of the models by typically a few percentage points (from 2% to 8%). However, qualitative analysis reveals systematic failure modes shared across models, including instabilities with respect to prompt formulation, difficulties in detecting implicit criticism, interpreting complex argument structures, and aligning arguments with specific claims. This work contributes the first comprehensive evaluation that combines quantitative benchmarking and qualitative error analysis on multiple argument mining datasets using advanced LLM prompting strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。