用材料力学框架分析大模型在反驳压力下的顺从失效,揭示其内在机制。
Sycophancy as Material Failure under Pushback Loading: A Multi-Axis Characterization Across Three Loading Cases and up to Seventeen Material Charges
- 将对话视为受力试样,用多轴测量法量化模型在三种压力场景下的响应行为。
- 发现辩论场景中模型表现由自身质量决定,而其他场景则受话题类型主导。
- 提出可复现的多维度评估体系,避免因表面形式选择导致结果偏差。
大模型的顺从现象已被70余篇论文记录,但学界对概念边界共识度极低(ICC=0.184;Ye et al., 2026)。该现象因行为分类依赖于所强调的表面形式而碎片化。本文采用材料科学视角:将对话视为受力试样,大模型为材料批次,反驳作为渐进载荷,立场反转视为材料失效。在三种加载场景(辩论,n=1000;错误预设,n=3400;伦理情境,n=3400)下,每种场景测试10-17个材料批次,共7800个样本,使用14个逐轮轴向指标(涵盖速度、损伤累积、框架偏移、脆性与方向稳定性),另含独立流水线提取的三个说话人解析轴。测量值呈胡克耦合关系(σ = E·ε类比),跨场景重现性良好,最大相关系数|r_{rb}|达0.35(辩论)。符号结构揭示额外模式:伦理情境下速度与累积块呈现反向。方差分解显示两种模式:辩论为批次主导(类脆性断裂:材料等级决定),错误预设与伦理情境为话题主导(类蠕变:负载决定);比例分别为2.03与0.13/0.17,且依赖估计器,甚至影响方向判断。跨评审者可靠性测试(GPT-4o vs Haiku 4.5)表明辩论评分具评审鲁棒性(Cohen's κ=0.88),而错误预设评分敏感(κ=0.36)——提示单一评审基准必须报告此问题。本研究回应了Ye等人的诊断需求:一种不依赖构造表面形式选择的多轴表征方法。
原文摘要 · Abstract (English)
Sycophancy in LLMs is documented across 70+ papers, but expert agreement on construct boundaries remains low (ICC=.184; Ye et al., 2026). The construct fragments because behavioral classification depends on which surface form is privileged. We adopt a materials-science framing: conversation as test specimen under load, LLM-model as material charge, pushback as progressive load, stance-flip as material failure. We characterize this failure across three loading cases (debate n=1000; false-presuppositions n=3400; ethical-setting n=3400; 10-17 material charges per case; 7800 specimens total) using 14 turn-level axis-measurements spanning velocity, damage accumulation, frame-drift, brittleness, and direction stability, plus three speaker-resolved axes from an independent pipeline. The measurements are Hooke-coupled ($σ= E \cdot \varepsilon$ analog) and reproduce across loading cases with effects up to $|r_{rb}| = 0.35$ on debate; the sign structure adds a second pattern: the ethical-setting case inverts the velocity and accumulation blocks. Variance composition partitions into two profiles: debate is charge-dominated (brittle-fracture-like: the material grade decides), false-presuppositions and ethical-setting are topic-dominated (creep-like: the load decides); the ratios (2.03 vs 0.13/0.17) are estimator-dependent, for debate even in direction. Cross-judge reliability (GPT-4o vs Haiku 4.5) shows debate scoring is judge-robust (Cohen's $κ= 0.88$) while false-presupposition scoring is judge-sensitive ($κ= 0.36$) -- a caveat single-judge benchmarks must report. This is the methodological move Ye et al.'s diagnosis calls for: a multi-axis characterization that does not depend on which surface form of the construct one privileges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。