提出让大模型真正理解情境的对齐方法,解决表面流畅但实际脆弱的问题。
Beyond Surface Alignment: Grounding the Dynamics of Situational Understanding and Generative Control in LLMs

- 通过分析输入处理与生成结构,建立情境与生成的深层对齐机制
- 发现主流模型在长上下文下难以维持一致心智模型,生成易陷入风格崩溃
- 适用于高风险场景如戒瘾支持,推动生成式智能向真实情境扎根
当前大语言模型的对齐调优侧重表面行为——流畅性、安全性与语调一致性。尽管适合日常对话,但本文指出这种表面对齐掩盖了缺乏根基的问题,导致模型虽风格自信却情境脆弱。我们提出「有根基对齐」框架,分析模型如何处理上下文(输入)与生成结构(输出),并将其与人类需求对齐。首先,通过SitTest发现,即便具备大上下文窗口,前沿模型仍难以维持动态环境的一致心智模型;ReCode进一步显示,模型依赖表面启发式而非深层句法依赖,即“读取”大量历史却未真正“理解”演变情境。其次,评估生成根基性:引入分支因子(BF)映射生成路径,发现标准对齐调优使生成空间过早风格坍缩;回溯测试(Hindsight)表明,模型常无法理解自身生成内容。最后,提出动态控制机制:AI Realtor实现上下文工程以弥补情境缺失;基线对齐模型协作将探索与风格约束解耦;采用渐进采样实现可验证强化学习,并应用于戒瘾支持场景,模型生成的合理化内容为高风险领域提供沟通接口。整体工作推动模型从表面对齐迈向基于上下文与生成的深层锚定。
原文摘要 · Abstract (English)
The current alignment tuning paradigm for Large Language Models (LLMs) prioritizes surface-level behaviors -- fluency, safety, and tonal consistency. While effective for casual chat, this thesis argues that such surface alignment masks a lack of grounding, creating models that are stylistically confident but situationally brittle. We propose a framework of Grounded Alignment, analyzing how models process context (Input) and structure generation (Output), then aligning these grounded behaviors to human needs. First, we evaluate failures in Situational Grounding. SitTest shows that despite large context windows, state-of-the-art models struggle to maintain a consistent "mental model" of a changing environment. ReCode further shows that models rely on surface heuristics rather than deep syntactic dependencies: they "read" extensive histories without truly "understanding" the evolving situation. Second, we evaluate Generative Grounding. We introduce the Branching Factor (BF) to map LLM generation, finding that standard alignment tuning constricts this landscape into premature stylistic collapse. Hindsight further shows that models often fail to understand their own generations. Finally, we propose Dynamic Control for grounded interaction. AI Realtor demonstrates context engineering to compensate for poor situational grounding. Base-Aligned Model Collaboration decouples exploration from stylistic constraints. We also present Annealed Sampling for verifiable reinforcement learning and apply these ideas to Addiction Support, where model-generated rationalization offers a communication interface for high-stakes domains. Collectively, this work moves beyond surface alignment toward agents anchored in both context and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。