arXiv:2502.14501cs.CL2025-02NAACL被引 6

提出多视角评估论点质量的新方法,打破单一标准局限。

Towards a Perspectivist Turn in Argument Quality Assessment

  • 构建论点质量标注的多层分类体系,整合多种评估维度。
  • 发现现有数据集多缺乏个体标注者信息,制约多视角研究。
  • 建议未来研究采用可控标注者群体,提升评估多样性与可靠性。

论点质量评估依赖逻辑、修辞与辩证等属性,但这些属性本身具有主观性,允许多种合理判断共存,且无绝对真值。这与机器学习中接纳多元视角的趋势一致,但在自然语言处理领域尚未充分探索。核心障碍之一是缺乏合适的标注数据集。本文通过系统梳理现有论点质量数据集,建立双维分类体系:(a) 标注内容——汇总各数据集覆盖的质量维度,并构建统一分类体系以增强可比性与互操作性;(b) 标注者信息——调查数据集中关于标注者的描述,支持多视角研究。我们识别出适合构建多视角模型的数据集(即包含个体非聚合标注的数据),并在初步研究中展示控制标注者选择的重要性。

原文摘要 · Abstract (English)

The assessment of argument quality depends on well-established logical, rhetorical, and dialectical properties that are unavoidably subjective: multiple valid assessments may exist, there is no unequivocal ground truth. This aligns with recent paths in machine learning, which embrace the co-existence of different perspectives. However, this potential remains largely unexplored in NLP research on argument quality. One crucial reason seems to be the yet unexplored availability of suitable datasets. We fill this gap by conducting a systematic review of argument quality datasets. We assign them to a multi-layered categorization targeting two aspects: (a) What has been annotated: we collect the quality dimensions covered in datasets and consolidate them in an overarching taxonomy, increasing dataset comparability and interoperability. (b) Who annotated: we survey what information is given about annotators, enabling perspectivist research and grounding our recommendations for future actions. To this end, we discuss datasets suitable for developing perspectivist models (i.e., those containing individual, non-aggregated annotations), and we showcase the importance of a controlled selection of annotators in a pilot study.

论点评估多视角数据集主观性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。