用大模型分析网络评论中的预定义论点,效果不错但有局限。
LLMs for Argument Mining: Detection, Extraction, and Relationship Classification of pre-defined Arguments in Online Comments
- 采用四种顶尖大模型处理论点检测、抽取和关系分类任务。
- 在六类争议话题的2000+评论上表现良好,大模型效果更优。
- 对长句、复杂情绪语句处理不佳,适合研究者参考使用。
针对堕胎等争议议题的公开讨论进行大规模自动化分析,需要识别并理解其中使用的论点。尽管大语言模型(LLMs)在自然语言处理任务中表现出色,但其在挖掘在线评论中特定主题的预定义论点方面仍缺乏系统研究。本文在涵盖六类极化话题、超过2000条意见评论的数据集上,评估了四种先进LLMs在三个论点挖掘任务中的表现。定量结果表明,整体性能优异,尤其大型且经过微调的LLMs表现突出,但存在显著环境成本。详细错误分析揭示,在长句和语义复杂的评论、以及情感强烈的表达中存在系统性不足,这对内容审核或观点分析等下游应用构成潜在风险。研究结果凸显了LLMs在在线评论自动论点分析中的潜力与当前局限。
原文摘要 · Abstract (English)
Automated large-scale analysis of public discussions around contested issues like abortion requires detecting and understanding the use of arguments. While Large Language Models (LLMs) have shown promise in language processing tasks, their performance in mining topic-specific, pre-defined arguments in online comments remains underexplored. We evaluate four state-of-the-art LLMs on three argument mining tasks using datasets comprising over 2,000 opinion comments across six polarizing topics. Quantitative evaluation suggests an overall strong performance across the three tasks, especially for large and fine-tuned LLMs, albeit at a significant environmental cost. However, a detailed error analysis revealed systematic shortcomings on long and nuanced comments and emotionally charged language, raising concerns for downstream applications like content moderation or opinion analysis. Our results highlight both the promise and current limitations of LLMs for automated argument analysis in online comments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。