arXiv:2412.08795cs.CLcs.AI2024-12NAACL被引 3

提出新公平性度量,更准确评估多文档摘要中不同群体信息覆盖。

Coverage-based Fairness in Multi-document Summarization

  • 基于文档覆盖率设计新公平性度量,考虑输入文档冗余。
  • 发现几乎所有LLM都过度代表某些社会属性值,Claude3-sonnet最公平。
  • 适用于评估大模型摘要公平性,尤其关注社会属性代表性。

多文档摘要中的公平性衡量系统是否能公平地呈现具有不同社会属性值的文档信息。以往研究采用基于统计平等的按比例代表(Proportional Representation)来量化摘要层面的公平性,但该方法未考虑输入文档内的冗余,也忽略了语料库层面的不公平。本文提出一种新的摘要级公平性度量——等覆盖率(Equal Coverage),基于不同社会属性值文档的覆盖率,并考虑文档内部冗余。为检测语料库级不公平,我们还提出了覆盖率平等(Coverage Parity)这一新度量。人工评估显示,我们的度量更符合公平性定义。利用这些度量,我们评估了13个不同大语言模型的公平性,结果表明Claude3-sonnet在所有模型中最为公平;同时发现几乎所有模型均过度代表某些社会属性值。代码已公开于https://github.com/leehaoyuan/coverage_fairness。

原文摘要 · Abstract (English)

Fairness in multi-document summarization (MDS) measures whether a system can generate a summary fairly representing information from documents with different social attribute values. Fairness in MDS is crucial since a fair summary can offer readers a comprehensive view. Previous works focus on quantifying summary-level fairness using Proportional Representation, a fairness measure based on Statistical Parity. However, Proportional Representation does not consider redundancy in input documents and overlooks corpus-level unfairness. In this work, we propose a new summary-level fairness measure, Equal Coverage, which is based on coverage of documents with different social attribute values and considers the redundancy within documents. To detect the corpus-level unfairness, we propose a new corpus-level measure, Coverage Parity. Our human evaluations show that our measures align more with our definition of fairness. Using our measures, we evaluate the fairness of thirteen different LLMs. We find that Claude3-sonnet is the fairest among all evaluated LLMs. We also find that almost all LLMs overrepresent different social attribute values. The code is available at https://github.com/leehaoyuan/coverage_fairness.

摘要公平性大模型评估社会属性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。