为评估主观性模型提出7项关键标准,推动更公平的AI评价体系。
Not All Subjectivity Is the Same! Defining Desiderata for the Evaluation of Subjectivity in NLP
- 从用户视角出发构建主观性评估标准
- 发现60篇论文中主观性研究存在多重盲区
- 适合关注公平性与多元声音的NLP研究者
主观判断是多个自然语言处理数据集的重要组成部分,近期研究越来越重视能够反映多元视角的模型输出。这类回应有助于揭示少数群体的声音,这些声音常被主流观点掩盖。然而,当前的评估实践是否与模型目标一致仍存疑问。本文提出七项主观性敏感模型的评估理想标准,基于主观性在NLP数据与模型中的呈现方式。标准采用自上而下的方法设计,注重用户中心影响。通过对60篇论文实验设置的分析,发现主观性研究仍存在诸多不足:输入中模糊与多声部的区别未被区分、主观性是否有效传达给用户尚不明确,以及各项标准之间缺乏协同作用等。
原文摘要 · Abstract (English)
Subjective judgments are part of several NLP datasets and recent work is increasingly prioritizing models whose outputs reflect this diversity of perspectives. Such responses allow us to shed light on minority voices, which are frequently marginalized or obscured by dominant perspectives. It remains a question whether our evaluation practices align with these models' objectives. This position paper proposes seven evaluation desiderata for subjectivity-sensitive models, rooted in how subjectivity is represented in NLP data and models. The desiderata are constructed in a top-down approach, keeping in mind the user-centric impact of such models. We scan the experimental setup of 60 papers and show that various aspects of subjectivity are still understudied: the distinction between ambiguous and polyphonic input, whether subjectivity is effectively expressed to the user, and a lack of interplay between different desiderata, amongst other gaps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。