用小模型研究神经网络如何做道德判断,发现偏见集中在特定计算阶段。
Building Interpretable Models for Moral Decision-Making
- 设计两层变换器处理道德困境,用嵌入编码影响对象与数量
- 在Moral Machine数据上达到77%准确率,模型规模适合深度分析
- 通过可解释性技术发现道德推理分布在不同计算阶段
我们构建了一个定制的两层Transformer模型,用于研究神经网络在电车难题类情境下的道德决策机制。该模型通过嵌入向量编码受影响者身份、人数及结果归属等结构化信息。模型在Moral Machine数据集上达到77%的准确率,同时保持较小规模,便于深入分析。我们采用多种可解释性技术,揭示了道德推理在神经网络中分布于不同计算阶段,且偏见集中于特定阶段,其他发现也进一步验证了这一分布特性。
原文摘要 · Abstract (English)
We build a custom transformer model to study how neural networks make moral decisions on trolley-style dilemmas. The model processes structured scenarios using embeddings that encode who is affected, how many people, and which outcome they belong to. Our 2-layer architecture achieves 77% accuracy on Moral Machine data while remaining small enough for detailed analysis. We use different interpretability techniques to uncover how moral reasoning distributes across the network, demonstrating that biases localize to distinct computational stages among other findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。