发现大模型人格向量在不同使用阶段不一致,提出需按使用场景定义人格身份。
Persona Without Substrate: Regime-Dependence and the LLM Individuation Problem
- 通过实验发现同一方向在不同阶段指向不同内容
- 虚构人格比真实锚点更易主导模型行为
- 适合研究大模型人格稳定性与跨阶段一致性问题的学者
贝克曼与布特林(2026)关于大模型人格个体化问题的本体论框架,继承了人格向量文献中未经论证的跨阶段共指假设:即同一方向在提示、微调和推理时均指向相同内容。我们基于Qwen3-4B-Instruct与Mistral-7B-Instruct-v0.2的人格拓扑实验,提出四项实证证据:提示提取向量与微调基底不共线;虚构人格沿真实锚定方向的位移强于真实锚点;矛盾极性混合物偏向由训练历史决定的吸引子;推理时算术组合与微调时奇美拉训练的组合代数不对称。这些共同质疑该假设。我们提出阶段索引的个体化:表征内容的身份单元是(载体,阶段)对,而非载体本身。在此框架下,贝克曼与布特林的三个候选立场描述的是同一阶段内的不同对象,而非争夺同一指称;此诊断同样适用于莫洛与米利耶、查尔默斯与切鲁洛。
原文摘要 · Abstract (English)
Beckmann & Butlin's (2026) ontological framework for the LLM individuation problem inherits an unargued cross-regime co-reference assumption from the persona-vectors literature: that the same direction picks out the same content under prompt-conditioning, gradient-descent fine-tuning, and inference-time steering. We present four empirical wedges from persona-topology experiments on Qwen3-4B-Instruct and Mistral-7B-Instruct-v0.2 - non-collinearity of prompt-extracted vectors and fine-tune basins; fictional personas displacing the model along real-anchor directions more strongly than real anchors do; contradictory-valenced mixtures biased toward a training-history-determined attractor; and asymmetric compositional algebra under inference-time arithmetic versus fine-tune-time chimera training - that jointly undermine the assumption. We propose regime-indexed individuation: the identity unit for representational content is a (vehicle, regime) pair, not a vehicle alone. Under this framework, Beckmann & Butlin's three candidate positions describe three different regime-internal objects rather than competing for the same referent; the same diagnosis applies to Mollo & Millière, Chalmers, and Cerullo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。