arXiv:2602.12758eess.IVcs.AI2026-02

音频驱动人脸重建,低带宽下仍保持流畅视频会议。

VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction

  • 用音频驱动生成人脸视频,替代原始摄像头流。
  • 在32.80 kbps带宽下实现稳定传输,帧率与延迟显著改善。
  • 基于网络状态自动切换模式,适合弱网环境使用。

消费级和受限网络中的严重带宽不足可能破坏实时视频会议的稳定性:编码器速率管理趋于饱和,丢包率上升,帧率下降,端到端延迟显著增加。本文提出一种自适应会议系统,集成WebRTC媒体传输与补充的音频驱动说话头重建路径,并通过遥测数据驱动模式调控。系统包括WebSocket信令服务、可选的SFU多用户传输组件、支持实时提取WebRTC统计信息并导出CSV遥测数据的浏览器客户端,以及一个处理参考人脸图像和录音并生成合成MP4的AI REST服务;浏览器可将自身出站摄像头流替换为合成流,平均带宽仅32.80 kbps。系统还包含带宽模式切换策略和客户端模式状态日志记录功能。

原文摘要 · Abstract (English)

Intense bandwidth depletion within consumer and constrained networks has the potential to undermine the stability of real-time video conferencing: encoder rate management becomes saturated, packet loss escalates, frame rates deteriorate, and end-to-end latency significantly increases. This work delineates an adaptive conferencing system that integrates WebRTC media delivery with a supplementary audio-driven talking-head reconstruction pathway and telemetry-driven mode regulation. The system consists of a WebSocket signaling service, an optional SFU for multi-party transmission, a browser client capable of real-time WebRTC statistics extraction and CSV telemetry export, and an AI REST service that processes a reference face image and recorded audio to produce a synthesized MP4; the browser can substitute its outbound camera track with the synthesized stream with a median bandwidth of 32.80 kbps. The solution incorporates a bandwidth-mode switching strategy and a client-side mode-state logger.

视频会议低带宽语音驱动AI重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。