arXiv:2602.09070cs.SDcs.AI2026-02

用情绪编码叙事逻辑,自动生成长视频配乐。

NarraScore: Bridging Visual Narrative and Musical Dynamics via Hierarchical Affective Control

  • 以情绪轨迹替代复杂叙事分析,实现高效配乐生成。
  • 在10分钟以上视频上保持高一致性,计算开销极低。
  • 适合影视制作、短视频创作等需要自动配乐的场景。

为长视频合成连贯配乐仍面临三大挑战:计算可扩展性、时间一致性以及对叙事逻辑演变的语义盲视。为此,我们提出NarraScore,一种基于情感作为叙事逻辑高密度压缩的核心思想的分层框架。独特之处在于,将冻结的视觉-语言模型(VLMs)重用于连续情感传感器,从高维视觉流中提取出具有叙事感知的效价-唤醒轨迹。机制上,NarraScore采用双分支注入策略,协调全局结构与局部动态:全局语义锚点确保风格稳定,而手术式令牌级情感适配器通过逐元素残差注入调节局部张力。该极简设计规避了密集注意力和架构克隆的瓶颈,有效缓解数据稀缺带来的过拟合风险。实验表明,NarraScore在保持卓越一致性和叙事对齐的同时,计算开销可忽略不计,建立了一个完全自主的长视频配乐生成范式。

原文摘要 · Abstract (English)

Synthesizing coherent soundtracks for long-form videos remains a formidable challenge, currently stalled by three critical impediments: computational scalability, temporal coherence, and, most critically, a pervasive semantic blindness to evolving narrative logic. To bridge these gaps, we propose NarraScore, a hierarchical framework predicated on the core insight that emotion serves as a high-density compression of narrative logic. Uniquely, we repurpose frozen Vision-Language Models (VLMs) as continuous affective sensors, distilling high-dimensional visual streams into dense, narrative-aware Valence-Arousal trajectories. Mechanistically, NarraScore employs a Dual-Branch Injection strategy to reconcile global structure with local dynamism: a \textit{Global Semantic Anchor} ensures stylistic stability, while a surgical \textit{Token-Level Affective Adapter} modulates local tension via direct element-wise residual injection. This minimalist design bypasses the bottlenecks of dense attention and architectural cloning, effectively mitigating the overfitting risks associated with data scarcity. Experiments demonstrate that NarraScore achieves state-of-the-art consistency and narrative alignment with negligible computational overhead, establishing a fully autonomous paradigm for long-video soundtrack generation.

音频生成叙事理解视觉语言模型情感建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。