arXiv:2605.26486cs.CV2026-05被引 4

开源视频人物生成框架,实现高保真语音驱动与长期稳定输出。

LongCat-Video-Avatar 1.5 Technical Report

论文配图:LongCat-Video-Avatar 1.5 Technical Report
图 1 · 摘自论文原文
  • 升级音频编码器至Whisper Large,优化训练流程提升稳定性。
  • 支持长视频生成,身份一致性强,多人群体与复杂动作表现自然。
  • 推理速度达8步,兼顾效率与画质,适合工业级部署应用。

尽管语音驱动视频生成技术取得进展,但实现商业级稳定性仍具挑战。我们提出LongCat-Video-Avatar 1.5,一个以系统工程和生产就绪性为导向的开源框架,超越架构创新。通过将音频编码器升级为Whisper Large,并精细化调整训练策略,v1.5实现了精准口型同步、全身时序稳定以及严格身份一致的长视频生成能力。经过严谨的数据清洗与基于人类反馈的强化学习(RLHF)训练,模型可泛化至动漫、动物等风格化领域,且原生支持多人互动与物体操作等复杂现实场景。针对工业部署需求,采用先进步数蒸馏技术,将推理步骤降至8个非均匀步数(NFE),在服务效率与视觉质量间取得良好平衡。大量定量指标与涵盖500多个测试案例的严格人工评估验证了其优越性:在人类相似度评分与专家级质量评估中,v1.5表现优于或媲美主流闭源系统(如HeyGen、OmniHuman 1.5、Kling Avatar 2.0)。该开源发布显著缩小了学术原型与商用部署之间的差距。

原文摘要 · Abstract (English)

Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upgraded open-source framework prioritizing systematic engineering and production-readiness over architectural novelty. By upgrading the audio encoder to Whisper Large and meticulously scaling our training recipes, v1.5 achieves accurate lip-synchronization, full-body temporal stability, and robust long-video generation with strict identity consistency. Through rigorous data curation and RLHF Training, the model readily generalizes to stylized domains such as anime and animals, and natively handles complex real-world conditions, such as multi-person interactions and object handling. Furthermore, addressing the practical demands of industrial deployment, we employ advanced step distillation to accelerate inference to an optimal 8 NFE, achieving a favorable trade-off between serving efficiency and visual fidelity. The superiority of our approach is validated through extensive quantitative metrics and a rigorous human evaluation conducted on a comprehensive benchmark of over 500 diverse test cases. Results show that v1.5 achieves competitive or superior performance compared to leading closed-source systems (e.g., HeyGen, OmniHuman 1.5, Kling Avatar 2.0) across human-likeness ratings and expert-level quality assessments on our benchmark. With its open-source release, LongCat-Video-Avatar 1.5 narrows the gap between academic research prototypes and commercial-grade deployment.

视频生成语音驱动开源框架长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。