用多智能体框架提升文本生成视频中的文化准确性。
When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

- 分维度拆解提示词,由专责智能体并行或串行处理
- 在243个文化相关提示上,跨文化生成质量提升显著
- 适合关注跨文化内容生成的研究者与创作者
文本到视频(T2V)生成在视觉保真度上快速进步,但对单个提示中忠实呈现多种文化的能力仍研究不足。我们提出MAVEN,一种多智能体提示优化框架,用于提升单一文化和跨文化T2V生成中的文化保真度。MAVEN将提示分解为人物、动作和地点三个维度,由专门智能体并行或顺序处理。为支持系统评估,我们构建了一个包含243个文化基础提示和972个对应视频的新基准,覆盖中、美、罗三类文化,三种动作类别,以及单一与跨文化场景。结合CLIP指标、视觉语言模型评判及视频质量测量的评估显示,多智能体优化,特别是并行专业化处理,显著提升了文化相关性,同时保持视觉质量与时间一致性。数据集与代码已公开于https://github.com/AIM-SCU/MAVEN。
原文摘要 · Abstract (English)
Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We introduce MAVEN, a multi-agent prompt refinement framework designed to improve cultural fidelity in both mono-cultural and cross-cultural T2V generation. MAVEN decomposes prompts into person, action, and location dimensions, handled by specialized agents operating in parallel or sequentially. To support systematic evaluation, we contribute a new benchmark of 243 culturally grounded prompts and 972 corresponding videos, spanning three cultures (Chinese, American, Romanian), three action categories, and both mono-cultural and cross-cultural scenarios. Evaluations combining CLIP-based metrics, VLM-as-judge assessments, and videoquality measures show that multi-agent refinement, particularly parallel specialization, significantly improves cultural relevance while preserving visual quality and temporal consistency. The dataset and code are available at https://github.com/AIM-SCU/MAVEN
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。