用强化学习教大模型按叙事理论讲逻辑故事
Retell, Reward, Repeat: Reinforcement Learning for Narrative Theory-Informed Story Retelling
- 结合结构主义叙事学与可度量叙事性,设计奖励机制
- 在时间旅行数据集上提升故事逻辑性、合理性和完整性
- 仅需少量查询数据,适合缺乏有效训练方法的叙事生成
反事实故事重述暴露了大模型在受限叙事空间中的不足,传统基于真实输出的微调无法教会模型生成合乎逻辑的叙事事件。本文提出Retell, Reward, Repeat(RRR)框架,融合结构主义叙事学与标量叙事性,构建基于强化学习的叙事生成训练管道。通过在TimeTravel数据集上引入人工标注的叙事平衡阶段,评估奖励模型表现。采用d-RLAIF方法,从文本特征中提取叙事性作为训练信号,无需参考输出。实验表明,RRR训练的大模型在逻辑性、合理性和完整性上均优于少样本和SFT基线,盲评验证输出质量。该方法仅依赖小规模查询数据,提供一种语言学基础扎实、成本可控的叙事生成后训练机制,凸显经典语言学理论在现代NLP中的持续价值。
原文摘要 · Abstract (English)
Counterfactual story retelling exposes LLM shortcomings in constrained narrative solution spaces where they can no longer rely on recalling memorised training data. Ground-truth-based post-training, such as SFT, fails to teach LLMs how to generate logical and rational narrative events. In this paper, we introduce Retell, Reward, Repeat (RRR), an RL-based pipeline synthesising Structuralist Narratology with scalar narrativity to teach storytelling structure. We extend the TimeTravel dataset with human-annotated stages of narrative equilibrium to evaluate reward models. By using d-RLAIF, RRR derives training signals from the narrativity of textual features without the need for reference outputs. Evaluations demonstrate that RRR-trained LLMs outperform few-shot and SFT baselines in logic, rationality, and completeness, with output quality additionally validated by blind human preference. Relying on a small, query-only dataset, RRR provides a linguistically grounded, cost-effective post-training mechanism for storytelling--a domain currently lacking effective post-training methods. RRR highlights the continued relevance of integrating established linguistic theories into contemporary NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。