梳理NLP中多样性量化的方法与意义,统一认知框架
A survey of diversity quantification in natural language processing: The why, what, where and how
- 基于生态与经济理论构建三维度多样性框架:多样性、均衡性、差异性
- 分析300+篇论文,揭示多样性在数据集与系统中的测量方法与场景
- 提出可比性强的评估体系,适合研究者规范多样性指标使用
近年来,自然语言处理(NLP)领域对多样性的关注持续上升,其被视为数据集与系统的重要属性,并已有多种度量方法。然而,多样性常被随意应用,缺乏明确依据,且跨研究间存在诸多不一致。本文受其他科学领域启发,借鉴Stirling(2007)提出的统一框架,该框架源自生态学与经济学,区分多样性三维度:品种多样性(variety)、均衡性(balance)和差异性(disparity)。我们系统调研了超过300篇来自ACL Anthology的多样性相关论文,构建了适用于NLP的四维分析框架:为何重要、测什么、在哪测、如何测。该分析提升了不同研究间多样性度量的可比性,揭示了新兴趋势,并为领域发展提供实践建议。
原文摘要 · Abstract (English)
The concept of diversity has received increasing attention in natural language processing (NLP) in recent years. It became an advocated property of datasets and systems, and many measures are used to quantify it. However, it is often addressed in an ad hoc manner, with few explicit justifications of its endorsement and many cross-paper inconsistencies. There have been very few attempts to take a step back and understand the conceptualization of diversity in NLP. To address this fragmentation, we take inspiration from other scientific fields where the concept of diversity has been more thoroughly conceptualized. We build upon Stirling (2007), a unified framework adapted from ecology and economics, which distinguishes three dimensions of diversity: variety, balance, and disparity. We survey over 300 recent diversity-related papers from ACL Anthology and build an NLP-specific framework with 4 perspectives: why diversity is important, what diversity is measured on, where it is measured, and how. Our analysis increases comparability of approaches to diversity in NLP, reveals emerging trends and allows us to formulate recommendations for the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。