arXiv:2605.25977cs.CLcs.AI2026-05

用100条专家思维链,让小模型学会创作质量对齐。

Creative Quality Alignment: Expert Tacit Knowledge Transfer via Chain-of-Thought Fine-Tuning

  • 仅用约100条专家思维链标注数据,实现创作质量对齐。
  • 发现现有对齐数据偏重手工技能,忽视受众与现实逻辑。
  • 揭示生成与评价的结构对偶性,解释小样本有效原因。

本文实现了Calibrated Surprise(Zou & Xu, 2026a)提出的创造质量度量的实证应用。研究核心问题是:该数学假设在工程层面是否成立?为使结论尽可能普适,我们采用最严苛的工程条件:低数据成本与小型基础模型。训练数据来自约100条由BC Protocol(Zou & Xu, 2026b)生成的专家思维链(CoT)注释。我们识别出一种数据偏差:当前公开对齐数据集普遍偏向工艺类知识,而受众建模与现实逻辑覆盖系统性薄弱。我们提出“创作质量对齐”(Creative Quality Alignment, CQA)这一工程方法类别。此外,提供一项支持性理论观察:在具有单一条件分布架构的大语言模型中,校准评价侧会通过架构对偶性自动传递至生成侧。这正是约100个思维链样本即可生效的结构性原因,而非如LIMA(Zhou et al., 2023)般的纯经验现象。

原文摘要 · Abstract (English)

This paper provides an empirical implementation of the creative quality metric proposed in Calibrated Surprise (Zou & Xu, 2026a). The question this paper addresses is: does this mathematical claim hold at the engineering level? To make the answer as general as possible, we deliberately choose the strictest engineering conditions: low data cost and a small base model. Training data comes from approximately 100 expert chain-of-thought (CoT) annotations produced by the BC Protocol (Zou & Xu, 2026b). We also identify a data bias: most publicly available alignment datasets are skewed toward craft-related knowledge, while audience modeling and reality-logic coverage are systematically weak. We use the term Creative Quality Alignment (CQA) to describe this class of engineering methods. We also offer a supporting theoretical observation: in an LLM with a single conditional distribution architecture, calibrating the appreciation side automatically transfers to the generation side via architectural duality. This is the structural reason why ~100 CoT examples are sufficient -- not a purely empirical observation like LIMA (Zhou et al., 2023).

创造质量思维链小样本对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。