合成数据带来新治理挑战,本文提出三项技术应对恶意行为、偏见和价值漂移。
Opportunities and Challenges of Frontier Data Governance With Synthetic Data
- 针对合成数据的三大风险,设计对应的技术机制。
- 在对抗训练、偏见缓解和价值强化中验证有效性。
- 适合关注生成式AI治理与安全的开发者与研究者。
合成数据(由机器学习模型生成)正成为解决数据获取难题的新方案,但其应用也带来了显著的治理与问责挑战,可能削弱现有的计算与数据治理范式。本文识别出合成数据引发的三大核心挑战:助长恶意行为、自发产生偏见以及价值观漂移。为此,我们提出了三项针对性技术机制,分别应用于对抗训练、偏见缓解与价值强化。这些机制不仅能有效应对合成数据带来的风险,更可作为未来前沿技术治理的关键抓手。
原文摘要 · Abstract (English)
Synthetic data, or data generated by machine learning models, is increasingly emerging as a solution to the data access problem. However, its use introduces significant governance and accountability challenges, and potentially debases existing governance paradigms, such as compute and data governance. In this paper, we identify 3 key governance and accountability challenges that synthetic data poses - it can enable the increased emergence of malicious actors, spontaneous biases and value drift. We thus craft 3 technical mechanisms to address these specific challenges, finding applications for synthetic data towards adversarial training, bias mitigation and value reinforcement. These could not only counteract the risks of synthetic data, but serve as critical levers for governance of the frontier in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。