arXiv:2607.10871cs.AIcs.HC2026-07中稿 · as an oral present…

用正念等理念构建可扩展的LLM心理健康对齐评估框架

Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health

  • 设计模块化流水线,支持模型、指标与基准灵活组合
  • 通过提示模块引入正念等伦理原则,提升对齐效果
  • 适合心理、伦理、AI交叉研究者使用

正念、慈悲等内省传统长期指导道德行为与利他互动,近期研究显示这些原则可能为大语言模型(LLMs)对齐提供有效范式,提升合作性并减少伦理违规。然而,随着新模型、评估指标和基准快速涌现,系统评估这些原则在多样且动态场景中的有效性仍具挑战,现有方法多为临时方案,难以泛化。本文提出一个模块化、可扩展的评估框架,初始聚焦心理健康领域,通过可复用的流水线实现新模型、指标与基准的无缝集成。该框架可复现现有最先进结果,并支持灵活混合搭配模型、指标与基准,实现公平对比与深层洞察。其即插即用的提示模块为引入正念等伦理视角提供了规范路径,使领域专家可在无需技术背景的前提下定义对齐标准。尽管初始聚焦心理健康,该框架具备领域无关性,可自然延伸至决策、道德推理与人机协作等领域。通过连接计算评估与以人为本的伦理推理,本工作为认知科学、行为经济学、哲学与系统设计的跨学科研究奠定基础,推动构建稳健、可信且有益社会的人机生态系统。

原文摘要 · Abstract (English)

Contemplative traditions have long guided ethical behavior and prosocial interaction, and recent work suggests that contemplative principles (e.g., mindfulness, compassion, non-dual reasoning) may offer a promising paradigm for aligning large language models (LLMs), improving cooperation and reducing ethical violations in LLM outputs. However, as new models, evaluation metrics, and benchmarks emerge rapidly, it remains challenging to systematically assess whether and how contemplative principles enhance LLM alignment across diverse and evolving scenarios, and existing approaches are often ad hoc and fail to generalize. We present a modular, extensible evaluation framework, initially targeted at the mental health domain, that enables seamless integration of new models, metrics, and benchmarks through a reusable pipeline. The framework currently reproduces existing state-of-the-art results and supports systematic cross-evaluation by flexibly mixing and matching models, metrics, and benchmarks, enabling fair comparison and deeper insight. Its plug-and-play prompting module offers a principled pathway for incorporating ethical perspectives such as contemplative principles, allowing domain experts to define alignment criteria without requiring technical expertise. Although initially focused on mental health, the framework is domain-agnostic and extends naturally to areas such as decision-making, moral reasoning, and human-AI collaboration. By bridging computational evaluation with human-centered ethical reasoning, this work lays the groundwork for interdisciplinary research spanning cognitive science, behavioral economics, philosophy, and system design, toward robust, trustworthy, and socially beneficial human-AI ecosystems.

LLM对齐心理健康伦理框架模块化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。