arXiv:2505.21819cs.CLcs.LG2025-05ICML被引 20

让生成模型按比例反映训练数据中的群体分布,解决多样性与偏见问题。

Representative Language Generation

  • 提出‘代表性生成’机制,要求输出比例匹配训练数据中各群体占比。
  • 定义‘群体闭包维数’作为关键组合量,分析可实现性与计算限制。
  • 证明在特定条件下可实现,但仅用成员查询无法计算,适合关注公平性的研究者。

我们引入‘代表性生成’,扩展Kleinberg等人(2024)提出的生成理论框架,并由Li等人(2024)形式化,以进一步解决生成模型中的多样性和偏见问题。该概念要求生成模型的输出能按比例代表训练数据中感兴趣的群体。我们刻画了代表性均匀与非均匀生成,引入‘群体闭包维数’作为关键组合量。针对极限情况下的代表性生成,我们分析了信息论与计算层面的问题,证明在某些条件下,对可数无限假设类和群体集合是可行的,但证明了仅使用成员查询无法实现可计算性。这与Kleinberg等人(2024)在标准生成极限下的正向结果形成对比。我们的成果为构建更多样化、更具代表性的生成模型提供了严格基础。

原文摘要 · Abstract (English)

We introduce "representative generation," extending the theoretical framework for generation proposed by Kleinberg et al. (2024) and formalized by Li et al. (2024), to additionally address diversity and bias concerns in generative models. Our notion requires outputs of a generative model to proportionally represent groups of interest from the training data. We characterize representative uniform and non-uniform generation, introducing the "group closure dimension" as a key combinatorial quantity. For representative generation in the limit, we analyze both information-theoretic and computational aspects, demonstrating feasibility for countably infinite hypothesis classes and collections of groups under certain conditions, but proving a negative result for computability using only membership queries. This contrasts with Kleinberg et al.'s (2024) positive results for standard generation in the limit. Our findings provide a rigorous foundation for developing more diverse and representative generative models.

生成模型公平性多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。