构建电商直播多模态理解模型,精准解析音视频与文本信息。
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

- 将音视频、图文等多模态数据统一映射到同一表征空间。
- 在直播长序列分析中实现时间对齐,提升关键信息定位精度。
- 适合需要实时多模态理解的电商场景,如直播带货智能问答。
电商直播需对噪声大、时序长的多模态流进行全模态理解,其中商品信息分布在语音、视频帧、产品图、叠加文字及用户提问中。我们提出TLive-Omni,一个专为直播电商设计的全模态理解模型,将图像、视频、音频和文本输入映射至统一表示空间。针对长时序直播分析,引入Per-vGrid机制,以时间标记的令牌组织方式将每个视频网格与其对应音频分组,并用显式边界令牌实现时间对齐。采用三阶段监督训练流程,逐步构建从全模态感知到指令响应的理解能力。进一步提出Faithful-RFT强化微调阶段,在满足实时性前提下,通过任务可验证反馈直接评分,提升回答忠实度与表达质量。此外,模型依托场景导向的原子能力分类体系与轻量级数据生成引擎,可将直播音视频流转化为语音识别、说话人分析、产品视觉定位、文本识别、时间定位、视频密集描述及全模态问答等多种训练信号。为支持大规模训练,采用同步长度分组采样减少填充,结合轻量动态采样策略,以近乎零奖励方差再生回放组,保持GRPO中相对优势的合理性。在电商直播基准测试中表现优异,同时在通用基准上也展现良好泛化能力。
原文摘要 · Abstract (English)
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries. We present TLive-Omni, an omni-modal understanding model tailored to live-commerce scenarios. It maps image, video, audio, and text inputs into a unified representation space. For long-form live streaming analysis, we introduce Per-vGrid, a timestamped token organization that groups each video grid with its temporally corresponding audio within explicit boundary tokens to facilitate temporal alignment. We design a three-stage supervised training recipe that progressively develops live-commerce understanding, from omni-modal perception to instruction-following responses. We then propose Faithful-RFT, a reinforcement fine-tuning stage that further improves answer faithfulness and expression quality while meeting real-time demands, scoring final responses directly with task-verifiable feedback rather than optimizing for reasoning-style exploration during rollout. Moreover, TLive-Omni is supported by a scenario-oriented atomic capability taxonomy and a compact data production engine that converts live-commerce audio, image, and video streams into training signals for speech recognition, speaker analysis, product visual grounding, text recognition, temporal grounding, video dense caption, and omni-modal QA, etc. For scalable training, a synchronized length-grouped sampler reduces padding while preserving comparable workloads across workers, while a lightweight dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful relative advantages for GRPO. Experiments on e-commerce live streaming benchmarks demonstrate strong performance across live-commerce domain tasks, together with excellent generalization on general benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。