递归生成文本数据时,性别偏见会趋向平衡而非持续放大。
Equilibrium Dynamics and Mitigation of Gender Bias in Synthetically Generated Data
- 通过三轮递归生成,发现偏见随初始水平向模型固有偏见收敛。
- 高初始偏见下降26%,低初始偏见上升36%,形成动态均衡。
- 性别对调增强法虽提升嵌入相似性偏见,却大幅降低下游任务偏见。
利用大语言模型进行递归提示可实现大规模合成数据生成,但可能加剧偏见。我们通过三种评估框架——基于规则的模式匹配、基于嵌入的语义相似性、下游任务表现——研究了三轮递归文本生成中的性别偏见动态。在三种初始偏见水平(0.1、0.3、0.6)和四种缓解策略下,结果表明偏见呈现均衡态而非单调放大:低初始偏见向模型固有偏见水平趋近(+36%),高初始偏见则衰减至该水平(-26%)。其中,引入性别对调样本的对比增强法,在低初始偏见下实现98.8%的下游偏见减少,平均达91%,尽管其嵌入层面偏见分数更高。这一矛盾揭示语义相似性指标与行为公平性可能脱节,强调合成数据生成中需多维度评估。
原文摘要 · Abstract (English)
Recursive prompting with large language models enables scalable synthetic dataset generation but introduces the risk of bias amplification. We investigate gender bias dynamics across three generations of recursive text generation using three complementary evaluation frameworks: rule-based pattern matching, embedding-based semantic similarity, and downstream task performance. Experiments with three initial bias levels (0.1, 0.3, 0.6) and four mitigation strategies reveal equilibrium dynamics rather than monotonic amplification. The low initial bias amplifies toward the model's inherent bias level (+36%), whereas the high initial bias decays toward it (-26%). Among mitigation methods, contrastive augmentation, which introduces gender-swapped variants, achieves significant downstream bias reduction (98.8% for low initial bias and 91% on average) despite producing higher embedding-based bias scores. This paradox demonstrates that semantic similarity metrics may diverge from behavioral fairness outcomes, highlighting the need for multidimensional evaluation in responsible synthetic data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。