arXiv:2503.00069cs.CYcs.AI2025-03被引 10

用社会框架提升大模型对齐,把模糊目标当机会而非缺陷。

Societal Alignment Frameworks Can Improve LLM Alignment

  • 引入社会、经济、契约等框架优化大模型对齐思路
  • 指出当前对齐困境源于目标不完整,而非技术不足
  • 提倡参与式界面设计,适合研究对齐与政策的学者

大语言模型(LLMs)的对齐研究聚焦于生成符合人类期望和共享价值观的响应。然而,由于人类价值观的复杂性与现有技术方法的局限性之间的根本脱节,对齐仍具挑战。当前方法常导致目标错设,反映出更广泛的不完全契约问题——即难以在开发者与模型间穷尽所有场景制定合同。本文主张,改进对齐需借鉴社会、经济与契约对齐框架,并探讨其潜在解决方案。鉴于社会对齐中固有的不确定性,我们进一步分析其在大模型对齐中的表现。最后,提出将目标未明确视为机遇而非缺陷的全新视角。除技术优化外,还强调参与式对齐界面设计的重要性。

原文摘要 · Abstract (English)

Recent progress in large language models (LLMs) has focused on producing responses that meet human expectations and align with shared values - a process coined alignment. However, aligning LLMs remains challenging due to the inherent disconnect between the complexity of human values and the narrow nature of the technological approaches designed to address them. Current alignment methods often lead to misspecified objectives, reflecting the broader issue of incomplete contracts, the impracticality of specifying a contract between a model developer, and the model that accounts for every scenario in LLM alignment. In this paper, we argue that improving LLM alignment requires incorporating insights from societal alignment frameworks, including social, economic, and contractual alignment, and discuss potential solutions drawn from these domains. Given the role of uncertainty within societal alignment frameworks, we then investigate how it manifests in LLM alignment. We end our discussion by offering an alternative view on LLM alignment, framing the underspecified nature of its objectives as an opportunity rather than perfect their specification. Beyond technical improvements in LLM alignment, we discuss the need for participatory alignment interface designs.

大模型对齐社会框架参与式设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。