arXiv:2510.20782cs.CLcs.AI2025-10

针对具体应用构建公平性评估数据集,提升大模型责任表现评测精准度。

A Use-Case Specific Dataset for Measuring Dimensions of Responsible Performance in LLM-generated Text

  • 基于真实商品描述任务,设计融合性别形容词与品类的标注提示集。
  • 发现模型在质量、真实性、安全性和公平性上存在多维度缺陷。
  • 适合关注负责任AI评估的研究者与工业界应用开发者。

当前大语言模型(LLM)评估多聚焦于文本生成等高层任务,未针对特定人工智能应用场景。该方法不足以衡量公平性等负责任AI维度,因在某一应用中关键的受保护属性在另一应用中可能不相关。本文构建了一个由实际应用驱动的数据集(生成给定产品特征的纯文本描述),通过将公平性属性与性别化形容词及产品类别交叉组合,生成丰富的带标签提示。我们展示了如何利用该数据识别LLM在质量、真实性、安全性与公平性方面的差距,并提出一种与具体用例匹配的评估框架,为研究社区提供可落地的资源。

原文摘要 · Abstract (English)

Current methods for evaluating large language models (LLMs) typically focus on high-level tasks such as text generation, without targeting a particular AI application. This approach is not sufficient for evaluating LLMs for Responsible AI dimensions like fairness, since protected attributes that are highly relevant in one application may be less relevant in another. In this work, we construct a dataset that is driven by a real-world application (generate a plain-text product description, given a list of product features), parameterized by fairness attributes intersected with gendered adjectives and product categories, yielding a rich set of labeled prompts. We show how to use the data to identify quality, veracity, safety, and fairness gaps in LLMs, contributing a proposal for LLM evaluation paired with a concrete resource for the research community.

负责任AI公平性评估数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。