大模型能模仿另一模型的幽默偏好,行为证据比指令更有效。
Cross-Model Humor Preference Modeling with Cards Against Humanity

- 用卡牌游戏测试模型间幽默偏好模仿能力,设计了五级渐进条件。
- 仅凭指令时准确率仅19%-26%,加入历史选择和理由后升至72.8%-82.3%。
- 证明模型可通过观察行为而非指令学习他人偏好,具备类心理理论能力。
本文研究一个大语言模型能否在受控的《人类之敌》风格任务中近似另一个模型的幽默偏好。以GPT-4o为裁判(Czar),Claude Opus-4.5为玩家(Player),构建二元幽默选择任务,确保成功不能来自自我偏好。通过反射单元稳定性程序,识别出244手具有确定但相反偏好的牌局,分为97手上下文池和147手留出测试集。玩家在五种渐进条件下评估:默认自我偏好、通用裁判建模指令、指定裁判模型、先前裁判选择、以及包含理由的先前选择。准确率从第1阶段0.7%提升至第2阶段19.0%和25.9%,再升至第5阶段72.8%和82.3%。综合检验确认每一步均显著提升。结果表明,角色指令与模型身份带来小幅改进,而行为证据——尤其是伴随理由——则促成显著跨模型偏好建模。该现象可解释为操作性而非表征性的类心理理论行为:玩家基于他者实际表现调整偏好,而非依赖内部心智状态表征。
原文摘要 · Abstract (English)
This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task. Two models - GPT-4o as Czar and Claude Opus-4.5 as Player - are evaluated on a binary humor-selection task constructed so that success cannot follow from self-preference. A reflected-cell stability procedure isolates 244 hands on which the two models hold deterministic but opposite preferences, partitioned into a 97-hand context pool and a 147-hand held-out test pool. The Player is then evaluated across five graded conditions: default self-preference, generic Czar-modeling instruction, model-identified Czar, prior Czar selections, and prior Czar selections with rationales. This gradient is designed to separate two sources of improvement: framing effects, in which the Player is told to attend to a Czar without seeing any of the Czar's behavior, and direct behavioral evidence, in which the Player is shown the Czar's prior choices. Player accuracy increased from 0.7% in Condition 1 to 19.0% and 25.9% in the framing-only conditions, and then rose to 72.8% and 82.3% once behavioral evidence and rationales were provided. An omnibus Cochran's Q test and pairwise McNemar tests confirmed that each step in the gradient produced a significant improvement. The results indicate that role instruction and model identity yield only modest gains, while behavioral evidence - especially when accompanied by rationales - supports substantial cross-model preference modeling. The findings are interpreted as theory-of-mind-like behavior in an operational rather than representational sense: the Player shifts away from self-preference toward another agent's demonstrated preferences, without any claim about an underlying representation of mental states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。