arXiv:2607.28384cs.AI2026-07

提出对齐框架,量化大模型在冲突指令下的偏好选择。

When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences

论文配图:When Specifications Conflict: A Symmetry-Based Framework for Measuring LLM Preferences
图 1 · 摘自论文原文
  • 构建对称冲突实验,直接观测模型决策偏好。
  • 发现形式化表达优于自然语言,输入输出示例最弱。
  • 适用于数学、代码、临床等多领域冲突分析。

大语言模型日益需整合存在矛盾的信息源,但缺乏可控且可解释的方法来分析其冲突解决机制。本文提出一种受控实验框架,通过构造具有明确冲突的规范,使模型在竞争规范间的抉择得以直接观测与分析。基于对称性设计,该框架减少混杂因素,实现不同表示类型间偏好的系统比较。在包含550个冲突实例、覆盖11类函数的可执行数学基准上评估,对比纯自然语言、形式语言、自然化形式语言及输入-输出示例四种表示类型。结果显示模型表现出系统性偏好而非随机行为,其偏好顺序为:形式语言 ≈ 自然化形式语言 > 纯自然语言 > 输入-输出示例。示例效应还受模型能力与函数族影响。框架进一步扩展至布尔代数、代码生成和临床领域中的异构规范冲突,验证其跨任务、跨形式的普适性。该框架提供了一种统一测量大模型处理信息冲突的能力的方法。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly required to integrate multiple sources of information that may be inconsistent or conflicting. However, there is still a lack of controllable and attributable methods for analyzing how models resolve conflicts between competing specifications. We propose a controlled experimental framework for studying model preferences under conflicting specifications. By constructing specifications with explicit conflicts, the framework enables model choices between competing specifications to be directly observed and analyzed. A symmetry-based design further reduces confounding factors, allowing preferences across representation types to be compared systematically. We evaluate the framework on an executable mathematical benchmark with 550 conflict instances spanning 11 function families, comparing four representation types: pure natural language, formal language, naturalized formal language, and input--output examples. Results show systematic preference patterns rather than random behavior, with a consistent ordering: $ \text{Formal} \approx \text{Naturalized Formal} > \text{Pure Natural Language} > \text{Input--Output Examples} $. Example effects further depend on model capability and function family. We extend the framework to heterogeneous specification conflicts in Boolean algebra, code generation, and the clinical domain, demonstrating its applicability across diverse tasks and specification forms. The framework provides a unified approach for measuring how LLMs resolve conflicts between competing sources of information.

大模型偏好冲突解决形式化表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。