arXiv:2605.11483cs.CL2026-05

用300条高质量文本让小模型学会斯多葛哲学内省美德。

StoicLLM: Preference Optimization for Philosophical Alignment in Small Language Models

  • 用偏好优化方法在极小数据集上训练小模型理解斯多葛哲学。
  • 仅需300条样本即可实现接近少样本提示的内省美德对齐效果。
  • 小模型无法习得斯多葛的普世责任,暴露其表征局限性。

大型语言模型虽擅长事实适应,但在严重数据约束下内化复杂哲学框架的能力仍待探索。本文通过在微小的斯多葛经典文本数据集上使用偏好优化(ORPO、AlphaPO)来专门化小型语言模型。经多模型批评者银行评估,仅需300条高保真样例即可使模型牢固对齐内在斯多葛美德,接近少样本提示效果并释放上下文窗口。然而,所有模型(包括少样本基线)均持续在斯多葛的外向性普世义务上表现失败,揭示小模型仅靠微数据集适配无法克服的表征局限。

原文摘要 · Abstract (English)

While large language models excel at factual adaptation, their ability to internalize nuanced philosophical frameworks under severe data constraints remains underexplored. We investigate this by specializing small LLMs on micro-datasets of foundational Stoic texts using preference optimization (ORPO, AlphaPO). Evaluated via a multi-model critic bank, our results show that just 300 high-fidelity examples can induce strong alignment with inward-facing Stoic virtues, closely approaching few-shot prompting while freeing the context window. Critically, however, all models, including few-shot baselines, exhibit a persistent failure on Stoicism's outward-facing cosmopolitan duties, pointing to a representational limitation of small models that micro-dataset adaptation alone cannot overcome.

小模型哲学对齐偏好优化斯多葛主义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。