为大模型时代的人工道德主体设计可操作的评估标准
Black Box Deployed -- Functional Criteria for Artificial Moral Agents in the LLM Era
- 提出十项功能标准评估大模型的道德能力
- 用自动驾驶公交场景验证标准实用性
- 适合研究人工智能伦理与安全的学者
强大的但不透明的大语言模型(LLMs)要求重新审视用于评估人工道德主体(AMAs)的哲学标准。传统框架依赖于可解释架构的假设,而LLMs因随机输出和内部状态不透明,无法满足这一前提。本文主张,传统伦理标准在实践中已过时。基于技术哲学核心议题,提出十项新功能标准:道德一致性、情境敏感性、规范完整性、元伦理意识、系统韧性、可信度、可纠正性、部分透明性、功能自主性与道德想象力。这些标准应用于“通过大语言系统模拟道德代理”(SMA-LLS),旨在引导基于LLM的道德主体在未来实现更好对齐与社会融合。通过涉及自主公共巴士(APB)的假设场景,展示其在高道德敏感情境中的实际应用价值。
原文摘要 · Abstract (English)
The advancement of powerful yet opaque large language models (LLMs) necessitates a fundamental revision of the philosophical criteria used to evaluate artificial moral agents (AMAs). Pre-LLM frameworks often relied on the assumption of transparent architectures, which LLMs defy due to their stochastic outputs and opaque internal states. This paper argues that traditional ethical criteria are pragmatically obsolete for LLMs due to this mismatch. Engaging with core themes in the philosophy of technology, this paper proffers a revised set of ten functional criteria to evaluate LLM-based artificial moral agents: moral concordance, context sensitivity, normative integrity, metaethical awareness, system resilience, trustworthiness, corrigibility, partial transparency, functional autonomy, and moral imagination. These guideposts, applied to what we term "SMA-LLS" (Simulating Moral Agency through Large Language Systems), aim to steer AMAs toward greater alignment and beneficial societal integration in the coming years. We illustrate these criteria using hypothetical scenarios involving an autonomous public bus (APB) to demonstrate their practical applicability in morally salient contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。