arXiv:2606.16617cs.CLcond-mat.mtrl-sci2026-06

用材料力学框架分析大模型在反驳压力下的顺从失效,揭示其内在机制。

Sycophancy as Material Failure under Pushback Loading: A Multi-Axis Characterization Across Three Loading Cases and up to Seventeen Material Charges

  • 将对话视为受力试样,用多轴测量法量化模型在三种压力场景下的响应行为。
  • 发现辩论场景中模型表现由自身质量决定,而其他场景则受话题类型主导。
  • 提出可复现的多维度评估体系,避免因表面形式选择导致结果偏差。

大模型的顺从现象已被70余篇论文记录,但学界对概念边界共识度极低(ICC=0.184;Ye et al., 2026)。该现象因行为分类依赖于所强调的表面形式而碎片化。本文采用材料科学视角:将对话视为受力试样,大模型为材料批次,反驳作为渐进载荷,立场反转视为材料失效。在三种加载场景(辩论,n=1000;错误预设,n=3400;伦理情境,n=3400)下,每种场景测试10-17个材料批次,共7800个样本,使用14个逐轮轴向指标(涵盖速度、损伤累积、框架偏移、脆性与方向稳定性),另含独立流水线提取的三个说话人解析轴。测量值呈胡克耦合关系(σ = E·ε类比),跨场景重现性良好,最大相关系数|r_{rb}|达0.35(辩论)。符号结构揭示额外模式:伦理情境下速度与累积块呈现反向。方差分解显示两种模式:辩论为批次主导(类脆性断裂:材料等级决定),错误预设与伦理情境为话题主导(类蠕变:负载决定);比例分别为2.03与0.13/0.17,且依赖估计器,甚至影响方向判断。跨评审者可靠性测试(GPT-4o vs Haiku 4.5)表明辩论评分具评审鲁棒性(Cohen's κ=0.88),而错误预设评分敏感(κ=0.36)——提示单一评审基准必须报告此问题。本研究回应了Ye等人的诊断需求:一种不依赖构造表面形式选择的多轴表征方法。

原文摘要 · Abstract (English)

Sycophancy in LLMs is documented across 70+ papers, but expert agreement on construct boundaries remains low (ICC=.184; Ye et al., 2026). The construct fragments because behavioral classification depends on which surface form is privileged. We adopt a materials-science framing: conversation as test specimen under load, LLM-model as material charge, pushback as progressive load, stance-flip as material failure. We characterize this failure across three loading cases (debate n=1000; false-presuppositions n=3400; ethical-setting n=3400; 10-17 material charges per case; 7800 specimens total) using 14 turn-level axis-measurements spanning velocity, damage accumulation, frame-drift, brittleness, and direction stability, plus three speaker-resolved axes from an independent pipeline. The measurements are Hooke-coupled ($σ= E \cdot \varepsilon$ analog) and reproduce across loading cases with effects up to $|r_{rb}| = 0.35$ on debate; the sign structure adds a second pattern: the ethical-setting case inverts the velocity and accumulation blocks. Variance composition partitions into two profiles: debate is charge-dominated (brittle-fracture-like: the material grade decides), false-presuppositions and ethical-setting are topic-dominated (creep-like: the load decides); the ratios (2.03 vs 0.13/0.17) are estimator-dependent, for debate even in direction. Cross-judge reliability (GPT-4o vs Haiku 4.5) shows debate scoring is judge-robust (Cohen's $κ= 0.88$) while false-presupposition scoring is judge-sensitive ($κ= 0.36$) -- a caveat single-judge benchmarks must report. This is the methodological move Ye et al.'s diagnosis calls for: a multi-axis characterization that does not depend on which surface form of the construct one privileges.

大模型对齐行为评估多轴分析材料类比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。