arXiv:2512.03026cs.CLcs.AI2025-12被引 2

提出无需数据集的持续伦理评估框架,实时检测大模型道德一致性。

The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models

  • 构建闭环系统,自动生成并评估伦理场景,实现无监督持续审计。
  • 发现伦理与毒性呈强负相关(rET = -0.81),与响应延迟几乎无关。
  • 适用于对模型长期行为安全有要求的研究者和开发者。

大型语言模型的快速发展凸显了道德一致性的重要性,即在不同情境下保持合乎伦理的推理能力。现有对齐框架多依赖静态数据集和事后评估,难以捕捉伦理推理在跨情境或时间尺度上的演变。本文提出道德一致性流水线(MoCoP),一种无需数据集、闭环运行的框架,可持续评估和解析大模型的道德稳定性。该框架包含三个模块:词法完整性分析、语义风险估计与基于推理的判断建模,构成自维持架构,能自主生成、评估并优化伦理情景。在GPT-4-Turbo和DeepSeek上的实验表明,MoCoP有效捕捉长期伦理行为,发现伦理维度与毒性维度呈强负相关(相关系数rET = -0.81,p < 0.001),与响应延迟几乎无关(rEL ≈ 0)。结果表明,道德连贯性与语言安全性是模型行为中稳定且可解释的特征,而非短期波动。通过将伦理评估重构为动态、模型无关的道德内省形式,MoCoP为可扩展的持续审计提供了可复现基础,推动了自主智能系统中计算伦理学的发展。

原文摘要 · Abstract (English)

The rapid advancement and adaptability of Large Language Models (LLMs) highlight the need for moral consistency, the capacity to maintain ethically coherent reasoning across varied contexts. Existing alignment frameworks, structured approaches designed to align model behavior with human ethical and social norms, often rely on static datasets and post-hoc evaluations, offering limited insight into how ethical reasoning may evolve across different contexts or temporal scales. This study presents the Moral Consistency Pipeline (MoCoP), a dataset-free, closed-loop framework for continuously evaluating and interpreting the moral stability of LLMs. MoCoP combines three supporting layers: (i) lexical integrity analysis, (ii) semantic risk estimation, and (iii) reasoning-based judgment modeling within a self-sustaining architecture that autonomously generates, evaluates, and refines ethical scenarios without external supervision. Our empirical results on GPT-4-Turbo and DeepSeek suggest that MoCoP effectively captures longitudinal ethical behavior, revealing a strong inverse relationship between ethical and toxicity dimensions (correlation rET = -0.81, p value less than 0.001) and a near-zero association with response latency (correlation rEL approximately equal to 0). These findings demonstrate that moral coherence and linguistic safety tend to emerge as stable and interpretable characteristics of model behavior rather than short-term fluctuations. Furthermore, by reframing ethical evaluation as a dynamic, model-agnostic form of moral introspection, MoCoP offers a reproducible foundation for scalable, continuous auditing and advances the study of computational morality in autonomous AI systems.

大模型伦理评估持续监控AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。