用合成数据让大模型精准遵守特定规则,适应动态政策变化。
SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data

- 从规范文档自动生成带边界意识的对齐数据
- 多模型测试中规则遵守率显著提升,不牺牲通用能力
- 适合需频繁更新合规策略的应用场景
随着大语言模型在真实场景中的广泛应用,对齐不再依赖单一的安全或有用性标准,而是由提供商或应用特有的模型规范决定。这些规范通常冗长、结构化且频繁更新,但现有对齐流程缺乏将其系统化为训练信号的机制。本文提出规范对齐新范式,将开发者编写的规范作为主要对齐目标,而非抽象原则或静态基准。为此,我们设计SpecAlign框架,直接从规范文档合成对齐数据。该方法结合结构化规则标注、可控规范实例化及多智能体对抗数据生成,产出细粒度、边界敏感的偏好对,同时捕捉合规行为与有意义的违规情形。在多个规范和骨干模型上的实验表明,使用SpecAlign训练能持续提升规则遵守度,保持通用能力,避免过度保守。结果表明,以显式规范为基底进行对齐,可实现快速、精确、可扩展的模型行为适配,满足不断演进的政策需求。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no longer governed by a single universal notion of safety or helpfulness, but instead by provider- or application-specific model specifications. These specifications are typically long, structured, and frequently updated, yet existing alignment pipelines lack a systematic mechanism to operationalize them as training signals. In this paper, we propose specification-grounded alignment, a new alignment paradigm that treats provider-authored model specifications as the primary alignment target rather than abstract principles or static benchmarks. To instantiate this paradigm, we introduce SpecAlign, a framework that synthesizes alignment data directly from specification documents. SpecAlign combines structured rule annotation, controllable specification instantiation, and multi-agent adversarial data synthesis to generate fine-grained, boundary-aware preference pairs that capture both compliant behaviors and meaningful specification violations. Experiments across multiple model specifications and backbone models demonstrate that training with SpecAlign consistently improves rule compliance while preserving general capabilities and avoiding over-conservative behavior. These results suggest that grounding alignment in explicit model specifications enables rapid, precise, and scalable adaptation of LLM behavior to evolving policy requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。