提出新基准与模型,让AI能持续生成百集以上的连贯音频剧。
NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama

- 用结构化潜变量建模长期剧情演化,保持200集以上连贯性。
- 新模型在200集时仍保持0.84以上剧情准确率,计算量仅为旧方法1/4。
- 支持多语言创作,跨语言质量提升0.23分,专业作者更偏爱其可控性。
长篇连续音频剧(200至800集)是重要创意媒介,但前沿大语言模型在此任务中表现不佳。我们对21种模型(涵盖经典、微调、开放/闭源前沿及推理类)在统一的叙事结构指标上进行评测。所有闭源系统在剧情节拍F1值上饱和于[0.78, 0.81]区间,并在长程(h=200)时下降约-0.20。我们提出NarrativeWorldBench,一个包含九项叙事结构指标的开放基准,评估范围覆盖h ∈ {10, 20, 50, 100, 200},并支持四种印地语族语言(印地语、泰米尔语、泰卢固语、马拉地语)的跨语言评测。我们引入N-VSSM——一种基于Mamba-2主干的叙事变分状态空间模型,通过事件条件后验维持256维结构化潜世界状态,配合80亿参数解码器,在超过200集的长程中实现≥0.84的剧情节拍F1,计算开销仅为闭源系统四分之一。学习到的文化迁移函数使跨语言保真度提升+0.20至+0.23个李克特点。在涉及12位专业作者的组内研究中(共240次实验),N-VSSM在长弧一致性上被偏好71%的时间,可控性评分高出1.3个李克特点。
原文摘要 · Abstract (English)
Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large language models (LLMs) fail. We benchmark 21 models, spanning classical, fine-tuned, open-frontier, closed-frontier, and reasoning tiers, on a uniform set of structural narrative metrics. All closed-frontier systems saturate at a plot-beat F1 in the band [0.78, 0.81] and collapse by about -0.20 F1 at horizon h=200. We introduce NarrativeWorldBench, an open benchmark of nine narrative-structure metrics evaluated across horizons h in {10, 20, 50, 100, 200}, with cross-lingual evaluation across four Indic languages (Hindi, Tamil, Telugu, Marathi). We introduce N-VSSM, a Narrative Variational State-Space Model that maintains a structured 256-dimensional latent world state over more than 200 episodes via a Mamba-2 backbone with an event-conditioned posterior and an 8B decoder. N-VSSM holds plot-beat F1 >= 0.84 across all horizons at 4x lower compute than the closed-frontier band. A learned Cultural Transfer Function lifts cross-language fidelity by +0.20 to +0.23 Likert points. In a within-subjects writer study (n = 12 professional authors, 240 trials), N-VSSM is preferred over Claude Opus 4.5 on long-arc consistency 71% of the time and rated +1.3 Likert points higher on controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。