新评测框架发现大模型客观能力不等于主观表现力,且现有工具可能察觉不到差异。
Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm

- 自演化评测系统自动设计行为维度,通过抗作弊机制自我优化并停止迭代
- 在无真人金标准下构建可信评测,人类评分相关性达0.45,仍可稳定识别能力迁移
- 发现模型在情感克制等主观能力上反而随规模退化,开源数据与评测模板
传统基准测试适用于可验证任务(如数学、代码),但大模型最快速发展的应用集中在主观与人际交互场景:陪伴、情绪支持、咨询。此时默认的有效性检验——将指标与人工判断相关联——缺乏稳定锚点:评分者间一致性低,受标注者身份影响,难以复现,且存在长度偏差。因此我们无法回答核心问题:在客观基准上提升的能力是否能迁移到主观行为?我们的工具能否发现这种转移的缺失?本文构建了适用于该场景的评测仪器,并揭示前沿结果。贡献包括:第一,一种自演化评测工具,可自主选择并生成行为维度,在乘法抗作弊适应度函数下自我优化,直至停止改进;第二,基于三重证书的“信任构建”范式,无需人类金标准即可建立可信度,人类评分相关性约0.45;第三,观测到能力迁移可分离:在49个模型、8个家族、24个月跨度中,主观行为表现未随客观指标提升而同步上升——最显著案例为‘建议克制’(知道何时不应提供建议),是所有维度中的普遍最低项;在gpt-4.1→gpt-5升级中,该维度甚至逆向退化,而整体得分掩盖了这一现象,仅一条指令即恢复。温暖克制能力由模型代际决定,而非单纯规模、MoE宽度、推理预算或推理模式;开源模型的帕累托前沿在约10-80倍更低单次调用成本下,匹配闭源旗舰性能;四类评审员在保留对话数据上复现了评测标准。数据、代码、锁定的评分体系与评审提示将在发表后公开。
原文摘要 · Abstract (English)
Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling. There the default validity test, correlating a metric to human judgment, has no stable anchor: inter-rater agreement is low, structured by annotator identity, barely reproducible, and length-biased. So we cannot answer the question that matters: does capability that scales on objective benchmarks transfer to subjective behavior, and would our instruments even tell us if it did not? We build an instrument for this regime and report what it reveals at the frontier. We contribute, first, a self-evolving instrument that selects and then authors its own behavioral dimensions under a multiplicative anti-gaming fitness, self-halting when it stops improving; second, a trust-by-construction paradigm that earns belief through three certificates established without a human gold standard, where human raters saturate (rho ~ 0.45); and third, the finding it makes visible -- capability transfer is dissociable. Across 49 models, 8 families, and 24 months, subjective behaviors are where objective-benchmark scaling fails to carry over: the sharpest case, advice-restraint (knowing when not to give advice), is the frontier's universal-lowest dimension, and at gpt-4.1->gpt-5 it ran backwards while the aggregate score hid it -- a regression one instruction recovers. Warm restraint is moved by model generation, not by raw scale, MoE width, inference budget, or reasoning mode; the open-weight Pareto frontier matches closed flagships at ~10-80x lower per-call cost; and four judge families replicate the rubric on held-out human ESConv conversations. Data, code, the locked rubric, and judge prompts will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。