用300万条道德寓言训练小模型,无需大模型也能生成有伦理意义的叙事。
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
- 用80亿参数以下模型生成三百万条结构化寓言,每则含角色、冲突、寓意等六要素。
- 生成成本仅0.135美元/千条,可在消费级硬件上完成,质量媲美大型模型。
- 数据集开源可复现,适合研究价值对齐、教育AI和叙事智能等方向。
道德故事是传递价值观的可靠方式,但现代NLP缺乏大规模、结构化的叙事语料库。我们提出TF1-EN-3M,据知是首个开放的、由不超过80亿参数的指令微调模型生成的三百万条英语寓言数据集。每则故事遵循六要素模板(角色→特质→背景→冲突→解决→道德),通过组合式提示引擎生成,确保体裁一致并覆盖广泛主题。评估采用多模型家族的开源大模型评判员,从语法、创意、道德清晰度与模板符合度等方面评分,并结合无参考的多样性与可读性指标。在十种开源生成器中,一个8B参数的Llama-3变体表现最优,可在消费级硬件上以约0.135美元/千条的成本生成高分寓言。我们公开数据集、生成代码、评估脚本与完整元数据,支持完全复现与成本基准测试。该数据集为指令遵循、叙事智能、价值对齐及儿童友好型教育AI研究开辟新路径,证明大规模道德叙事无需依赖专有大模型或封闭评估体系。
原文摘要 · Abstract (English)
Moral stories are a time-tested vehicle for transmitting values, yet modern NLP lacks a large, structured corpus that couples coherent narratives with explicit ethical lessons. We present TF1-EN-3M, to our knowledge the first open dataset of three million English-language fables generated exclusively by instruction-tuned models no larger than 8B parameters. Each story follows a six-slot scaffold (character -> trait -> setting -> conflict -> resolution -> moral), produced through a combinatorial prompt engine that guarantees genre fidelity while covering a broad thematic space. A fully reproducible evaluation pipeline employs a panel of open-weight LLM judges from distinct model families, scoring grammar, creativity, moral clarity, and template adherence, complemented by reference-free diversity and readability metrics. Among ten open-weight generator candidates, an 8B-parameter Llama-3 variant delivers the best quality-cost trade-off, producing high-scoring fables on consumer hardware at approximately $0.135 per 1,000 fables. We release the dataset, generation code, evaluation scripts, and full metadata under a permissive license, enabling exact reproducibility and cost benchmarking. TF1-EN-3M opens avenues for research in instruction following, narrative intelligence, value alignment, and child-friendly educational AI -- demonstrating that large-scale moral storytelling requires neither proprietary giant models nor proprietary evaluation infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。