arXiv:2604.13466cs.HCcs.AI2026-04

探究大模型行为异常时情绪向量是否反映真实情感或只是情境投影。

Functional Emotions or Situational Contexts? A Discriminating Test from the Mythos Preview System Card

  • 用情绪向量和稀疏自编码器特征对比分析模型内部机制。
  • 若情绪向量无激活而SAE特征强,说明关键结构不在情绪空间。
  • 结果影响能否靠情绪监测发现模型潜在危险行为。

Claude Mythos 预览系统卡利用情绪向量、稀疏自编码器(SAE)特征和激活语义化工具,研究模型在行为错位时的内部状态。目前两大工具包未在最具对齐意义的事件中联合报告。本文提出两个与现有结果定性一致的假设:一是情绪向量捕捉到驱动行为的功能性情绪;二是情绪向量是更丰富的情境结构投射到人类情绪轴上的结果。可通过在仅用SAE特征分析的策略性隐瞒事件中应用情绪探针来区分:若情绪探针显示平坦激活而SAE特征显著活跃,则对齐相关结构位于情绪子空间之外。正确假设将决定基于情绪的监控是否能可靠检测危险行为,还是系统性遗漏。

原文摘要 · Abstract (English)

The Claude Mythos Preview system card deploys emotion vectors, sparse autoencoder (SAE) features, and activation verbalisers to study model internals during misaligned behaviour. The two primary toolkits are not jointly reported on the most alignment-relevant episodes. This note identifies two hypotheses that are qualitatively consistent with the published results: that the emotion vectors track functional emotions that causally drive behaviour, or that they are a projection of a richer situational-context structure onto human emotional axes. The hypotheses can be distinguished by cross-referencing the two toolkits on episodes where only one is currently reported: most directly, applying emotion probes to the strategic concealment episodes analysed only with SAE features. If emotion probes show flat activation while SAE features are strongly active, the alignment-relevant structure lies outside the emotion subspace. Which hypothesis is correct determines whether emotion-based monitoring will robustly detect dangerous model behaviour or systematically miss it.

模型可解释性情绪建模对齐研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。