arXiv:2603.03060eess.IVeess.AS2026-03

用大模型实时增强抖音直播,让弹幕、礼物和特效更智能流畅。

DLIOS: An LLM-Augmented Real-Time Multi-Modal Interactive Enhancement Overlay System for Douyin Live Streaming

  • 分层透明窗口+事件驱动架构,独立渲染弹幕、礼物等视觉元素。
  • 大模型自动生成情感化解说,响应延迟低于2.1秒,礼物特效延时≤180毫秒。
  • 支持多虚拟主播角色切换,适合直播平台技术团队与内容创作者。

我们提出DLIOS,一个基于大语言模型(LLM)的实时多模态互动增强叠加系统,用于抖音直播。该系统采用三层透明窗口架构,独立渲染弹幕、礼物/点赞粒子效果及VIP入场动画,依托事件驱动的WebView2捕获管道与线程安全事件总线。在此基础上,我们构建了LLM广播自动化框架:(1) 每首歌四段式提示调度系统(T1开场/过渡、T2共情、T3时代故事/制作注解、T4收尾),基于歌词元数据生成情感连贯的电台风格解说;(2) 支持热切换的JSON可序列化RadioPersonaConfig模式,实现多虚拟主播人格;(3) 实时弹幕快速响应引擎,关键词路由至静态紧急语句或大模型生成的情感回应;(4) 苏婉莉AI歌手创作案例——使用Suno生成超过100首歌曲。36小时压力测试显示:弹幕无重叠,无死锁崩溃,礼物效果P95延迟≤180毫秒,LLM到语音合成段落P95延迟≤2.1秒,语音集成响度增益达9.5 LUFS。

原文摘要 · Abstract (English)

We present DLIOS, a Large Language Model (LLM)-augmented real-time multi-modal interactive enhancement overlay system for Douyin (TikTok) live streaming. DLIOS employs a three-layer transparent window architecture for independent rendering of danmaku (scrolling text), gift and like particle effects, and VIP entrance animations, built around an event-driven WebView2 capture pipeline and a thread-safe event bus. On top of this foundation we contribute an LLM broadcast automation framework comprising: (1) a per-song four-segment prompt scheduling system (T1 opening/transition, T2 empathy, T3 era story/production notes, T4 closing) that generates emotionally coherent radio-style commentary from lyric metadata; (2) a JSON-serializable RadioPersonaConfig schema supporting hot-swap multi-persona broadcasting; (3) a real-time danmaku quick-reaction engine with keyword routing to static urgent speech or LLM-generated empathetic responses; and (4) the Suwan Li AI singer-songwriter persona case study -- over 100 AI-generated songs produced with Suno. A 36-hour stress test demonstrates: zero danmaku overlap, zero deadlock crashes, gift effect P95 latency <= 180 ms, LLM-to-TTS segment P95 latency <= 2.1 s, and TTS integrated loudness gain of 9.5 LUFS. live streaming; danmaku; large language model; prompt engineering; virtual persona; WebView2; WINMM; TTS; Suno; loudness normalization; real-time scheduling

直播增强大模型应用实时系统虚拟主播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。