从群体偏好中学习社会价值体系,实现更贴近真实社会的AI价值对齐。
Learning the Value Systems of Societies from Preferences
- 基于启发式深度聚类,从人类偏好中自动识别共享价值基底与多元价值系统。
- 在真实旅行决策数据上验证,能有效捕捉社会内部的价值多样性。
- 适合研究社会价值观、伦理对齐与群体智能的学者与开发者。
将AI系统与人类价值观及利益相关者的价值体系对齐,是实现伦理人工智能的关键。在价值感知的AI系统中,决策依赖于个体价值的显式计算表征(价值基底)及其聚合形成的价值体系。由于这些价值难以人工获取和校准,价值学习方法旨在通过观察人类行为示范,自动推导出代理主体的价值与价值体系的计算模型。然而,社会科学与人文学科文献指出,将社会价值体系视为不同群体价值体系的集合,而非个体价值体系的简单叠加,更具合理性。因此,本文正式提出了学习社会价值体系的问题,并提出一种基于启发式深度聚类的方法。该方法通过观察一组代理的定性价值偏好,学习社会共享的价值基底及一组代表该社会多样性的价值系统。我们在一个真实的旅行决策数据用例中评估了该方法,结果表明其能够有效揭示社会内部的价值差异与共性。
原文摘要 · Abstract (English)
Aligning AI systems with human values and the value-based preferences of various stakeholders (their value systems) is key in ethical AI. In value-aware AI systems, decision-making draws upon explicit computational representations of individual values (groundings) and their aggregation into value systems. As these are notoriously difficult to elicit and calibrate manually, value learning approaches aim to automatically derive computational models of an agent's values and value system from demonstrations of human behaviour. Nonetheless, social science and humanities literature suggest that it is more adequate to conceive the value system of a society as a set of value systems of different groups, rather than as the simple aggregation of individual value systems. Accordingly, here we formalize the problem of learning the value systems of societies and propose a method to address it based on heuristic deep clustering. The method learns socially shared value groundings and a set of diverse value systems representing a given society by observing qualitative value-based preferences from a sample of agents. We evaluate the proposal in a use case with real data about travelling decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。