用作者身份理论评估大模型风格个性化,发现现有方法远未达到真实人类水平。
Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization
- 基于作者身份验证理论构建可校准的评估指标
- 四种方法均低于跨作者基准线(0.484-0.508)
- 不同评估方式结论截然相反,凸显理论根基重要性
风格个性化——让大模型模仿特定个体写作风格,而非仅适应任务偏好——缺乏基于作者身份科学的评估。本文将评估建立在作者身份验证理论基础上,揭示了基准测试可衡量内容的转变。结合三种测量传统:训练好的作者身份验证模型LUAR、解耦特征匹配的LLM作为评判者,以及经典功能词文体学,我们在50位作者和1000次生成上评估了四种推理时个性化方法。理论基础指标(LUAR)提供了其他非理论指标无法实现的校准基线(人类上限0.756,跨作者下限0.626),使评分具有绝对意义。所有方法得分均低于此下限(0.484-0.508),暴露了先前未被察觉的作者风格差距。三种指标间成对相关性近乎为零(|r| < 0.07),表明无理论支撑的评估选择会决定结论——一个LLM评判者声称有明显优劣,而LUAR则未发现有意义差异。这些发现展示了理论-基准循环的实际运作:作者身份理论揭示了常规基准所忽视的评估缺陷。
原文摘要 · Abstract (English)
Stylistic personalization - making LLMs write in a specific individual's style, rather than merely adapting to task preferences - lacks evaluation grounded in authorship science. We show that grounding evaluation in authorship verification theory transforms what benchmarks can measure. Drawing on three measurement traditions - LUAR (a trained authorship verification model), an LLM-as-judge with decoupled trait matching, and classical function-word stylometrics - we evaluate four inference-time personalization methods across 50 authors and 1,000 generations. The theory-grounded metric (LUAR) provides what ad hoc alternatives cannot: calibrated baselines (human ceiling 0.756, cross-author floor 0.626) that give scores absolute meaning. All methods score below this floor (0.484-0.508), exposing an authorship gap invisible to uncalibrated metrics. The three metrics produce near-zero pairwise correlations (|r| < 0.07), confirming that without theoretical grounding, metric choice determines conclusions - an LLM judge declares a clear winner while LUAR finds no meaningful differentiation. These findings demonstrate the theory-benchmark cycle in action: authorship theory exposes evaluation failures that ad hoc benchmarks miss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。