通过分析错误一致性,评估AI与人类决策的对齐程度。
Measuring Error Alignment for Decision-Making Systems
- 用错误匹配度和类别级错误相似性衡量系统间行为对齐
- 新指标与表征对齐度相关,且在多领域表现稳定
- 适合关注AI伦理与价值对齐的研究者使用
随着人工智能系统在决策中扮演越来越关键的角色,其可信度与可靠性成为核心关切。由于现代AI系统规模庞大、结构复杂,难以直接解释,需借助替代方法建立信任并判断其是否与人类价值观对齐。我们提出,衡量AI与人类信息处理相似性的有效指标,或可达成此目标。现有表征对齐(RA)方法虽能反映内部状态相似性,但收集人类数据成本高、难度大;而行为对齐(BA)方法更易获取,但其敏感性和可靠性仍存疑。本文提出两项新行为对齐度量:误分类一致率(misclassification agreement),用于衡量两系统在相同样本上的错误一致性;类别级错误相似性(class-level error similarity),用于衡量错误分布的相似性。实验表明,这两项指标与RA度量高度相关,并为现有BA指标提供互补信息,在多个领域中表现出良好性能,为价值对齐研究提供了新范式。
原文摘要 · Abstract (English)
Given that AI systems are set to play a pivotal role in future decision-making processes, their trustworthiness and reliability are of critical concern. Due to their scale and complexity, modern AI systems resist direct interpretation, and alternative ways are needed to establish trust in those systems, and determine how well they align with human values. We argue that good measures of the information processing similarities between AI and humans, may be able to achieve these same ends. While Representational alignment (RA) approaches measure similarity between the internal states of two systems, the associated data can be expensive and difficult to collect for human systems. In contrast, Behavioural alignment (BA) comparisons are cheaper and easier, but questions remain as to their sensitivity and reliability. We propose two new behavioural alignment metrics misclassification agreement which measures the similarity between the errors of two systems on the same instances, and class-level error similarity which measures the similarity between the error distributions of two systems. We show that our metrics correlate well with RA metrics, and provide complementary information to another BA metric, within a range of domains, and set the scene for a new approach to value alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。