让大模型从文学文本中自动推断舞台布景与人物动线
Text-to-Stage: Spatial Layouts from Long-form Narratives
- 用戏剧结构启发的确定性评估框架评测空间推理能力
- 结合拒绝采样与可验证奖励的强化学习,提升布局合理性
- 适合需要场景还原的影视改编、游戏叙事等应用
本文研究语言模型从无结构文本中进行空间推理的能力,模拟人类认知并自动化下游媒体应用流程。具体聚焦于‘文本转剧目’任务:从缺乏显式空间、位置或关系线索的文本中推断舞台剧布局(场景、角色位置、移动轨迹和房间类型)。我们提出一种受戏剧创作启发的确定性评估体系,并设计了一套训练与推理方案,结合基于最佳N次采样的拒绝式监督微调与通过GRPO实现的可验证奖励强化学习。在纯文本的古典英语文学语料库上实验表明,该方法在角色归属、空间合理性及动作经济性等多个指标上均优于基线模型,且与大型语言模型作为评判者及人工主观偏好保持一致。
原文摘要 · Abstract (English)
In this work, we probe the ability of a language model to demonstrate spatial reasoning from unstructured text, mimicking human capabilities and automating a process that benefits many downstream media applications. Concretely, we study the narrative-to-play task: inferring stage-play layouts (scenes, speaker positions, movements, and room types) from text that lacks explicit spatial, positional, or relational cues. We then introduce a dramaturgy-inspired deterministic evaluation suite and, finally, a training and inference recipe that combines rejection SFT using Best-of-N sampling with RL from verifiable rewards via GRPO. Experiments on a text-only corpus of classical English literature demonstrate improvements over vanilla models across multiple metrics (character attribution, spatial plausibility, and movement economy), as well as alignment with an LLM-as-a-judge and subjective human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。