arXiv:2504.15552cs.AI2025-04被引 1

用AI多智能体系统自动生成秦腔戏本,融合剧本、画面和配音。

A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models

  • 分三阶段:写剧本、生成场景图、合成配音,各司其职
  • 秦腔《窦娥冤》测试中综合评分3.6,比单智能体高0.3分
  • 去除非关键模块会显著降分,证明协同设计有效

本文提出一种新型多智能体框架,实现秦腔戏曲从头到尾的自动化创作,融合大语言模型、视觉生成与语音合成技术。三个专用智能体按序协作:Agent1利用大语言模型生成符合文化语境、逻辑连贯的剧本;Agent2通过视觉生成模型渲染情境匹配的舞台画面;Agent3则采用文本转语音技术输出情感丰富、同步精准的演唱。以《窦娥冤》为例,系统在剧本忠实度、视觉一致性、语音准确度上分别获得专家评分3.8、3.5和3.8,综合得分3.6,较单智能体基线提升0.3分。消融实验显示,移除Agent2或Agent3分别导致评分下降0.4和0.5分,凸显模块化协同的价值。该工作展示了人工智能流水线在传统表演艺术保护与推广中的潜力,并指明未来可在跨模态对齐、情感表达深度及扩展更多剧种方面进行优化。

原文摘要 · Abstract (English)

This paper introduces a novel multi-Agent framework that automates the end to end production of Qinqiang opera by integrating Large Language Models , visual generation, and Text to Speech synthesis. Three specialized agents collaborate in sequence: Agent1 uses an LLM to craft coherent, culturally grounded scripts;Agent2 employs visual generation models to render contextually accurate stage scenes; and Agent3 leverages TTS to produce synchronized, emotionally expressive vocal performances. In a case study on Dou E Yuan, the system achieved expert ratings of 3.8 for script fidelity, 3.5 for visual coherence, and 3.8 for speech accuracy-culminating in an overall score of 3.6, a 0.3 point improvement over a Single Agent baseline. Ablation experiments demonstrate that removing Agent2 or Agent3 leads to drops of 0.4 and 0.5 points, respectively, underscoring the value of modular collaboration. This work showcases how AI driven pipelines can streamline and scale the preservation of traditional performing arts, and points toward future enhancements in cross modal alignment, richer emotional nuance, and support for additional opera genres.

秦腔多智能体AI生成戏曲

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。