首个实时音视频社交世界模型,支持亚秒级交互与超长生成。
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

- 构建首个实时音视频自回归模型,支持单卡47.5帧/秒生成
- 采用新型训练技术,实现高效稳定训练与长期生成不漂移
- 专为社交互动设计,适合下一代AI社交平台研发
随着全球视频内容越来越多地在社交平台用于互动社交,面向社交世界的视频生成模型至关重要却长期被忽视。本文提出社交世界模型的概念,并构建首个原型模型MaineCoon。该模型为首个真正意义上的实时音视频自回归模型,拥有220亿参数,可在单张GPU上实现高达47.5帧/秒的生成速率,支持亚秒级交互。为实现高效稳定训练,引入自重采样、跨模态表征对齐、领域感知偏好优化及强化在线策略蒸馏(ROPD)等新技术。设计首个支持千秒级甚至更长生成的智能体流式推理框架,通过智能体缓存管理与提示规划缓解生成漂移。本工作不仅建立了高质量、低延迟、长时序音视频自回归模型的新SOTA基准,也指明了下一代AI原生社交平台的范式转变方向。
原文摘要 · Abstract (English)
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but largely overlooked by previous studies. In this work, we define the position of social world models and build a prototype model as the first step towards this goal. While previous world models successfully simulate physical environments or gaming world exploration, they remain fundamentally detached from human-centric social dynamics. To bridge this gap as the first step to social world models, we present MaineCoon, the first real-time audio-visual autoregressive model that has 22B parameters and is capable of real-time streaming generation and sub-second interaction, with a record-breaking frame rate of up to 47.5 FPS, on a single GPU. To the best of our knowledge, MaineCoon is also the first real-time audio-visual generation model specifically optimized for social-interactive applications. To enable efficient and stable training, we introduce several novel techniques into MaineCoon, including self-resampling, cross-modal representation alignment, domain-aware preference optimization, and reinforced online-policy distillation (ROPD). We also design the first agentic streaming inference framework that supports thousand-second-scale or even longer generation while mitigating drift with agentic cache management and prompt planing. These innovations significantly accelerate training while optimizing real-time inference performance. We believe this work not only sets a new state-of-the-art (SOTA) performance benchmark for high-quality, low-latency, and long-horizon audio-visual autoregressive models, but also points out the paradigm shift desired for next-generation AI-native social platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。