arXiv:2501.12405cs.CYcs.AI2025-01被引 4

提出三维对齐框架,拓展大模型对齐的边界。

Scopes of Alignment

  • 从能力、时效性、受众三方面重构对齐维度
  • 强调模型需具备实用能力与情境适配性
  • 适合关注模型实际应用的开发者与研究者

当前人工智能对齐研究多聚焦于将大语言模型等基础模型与无上下文的通用价值(如有益性、无害性、诚实性)对齐。前沿模型提供商也致力于此。本文论证需超越这一局限,提出三个新对齐维度:第一,能力对齐——模型为实现其预期用途必须具备的知识、技能或行为;第二,时效性对齐——根据使用场景区分语义或事件性时效;第三,受众对齐——针对大众、公共、小群体或一对一交流场景。最后,本文用该框架定位若干突破现有对齐观念的技术与工作流。

原文摘要 · Abstract (English)

Much of the research focus on AI alignment seeks to align large language models and other foundation models to the context-less and generic values of helpfulness, harmlessness, and honesty. Frontier model providers also strive to align their models with these values. In this paper, we motivate why we need to move beyond such a limited conception and propose three dimensions for doing so. The first scope of alignment is competence: knowledge, skills, or behaviors the model must possess to be useful for its intended purpose. The second scope of alignment is transience: either semantic or episodic depending on the context of use. The third scope of alignment is audience: either mass, public, small-group, or dyadic. At the end of the paper, we use the proposed framework to position some technologies and workflows that go beyond prevailing notions of alignment.

AI对齐模型能力应用场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。