深海通话变视频:用声波传文字,再合成同步口型视频
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
- 潜艇内语音转文字,通过声呐发送至水面
- 实现低延迟、高质量的口型同步视频重建
- 适合深海探测、极限环境通信研究者
本文报告了2022年夏季对泰坦尼克号残骸进行深潜时开展的通信实验。深海水中无法使用无线电,通信依赖声呐信号。由于声呐带宽极低且需传递可读数据,深海任务中采用文本消息。本文提出一种系统:在潜艇内将语音转为文字,发送至水面后重构为说话人唇同步的合成视频。该系统于2022年夏季实际下潜至泰坦尼克号残骸时进行了测试,实现了可接受的延迟与良好质量。系统演示视频见:https://youtu.be/C4lyM86-5Ig
原文摘要 · Abstract (English)
In this paper, we report on communication experiments conducted in the summer of 2022 during a deep dive to the wreck of the Titanic. Radio transmission is not possible in deep sea water, and communication links rely on sonar signals. Due to the low bandwidth of sonar signals and the need to communicate readable data, text messaging is used in deep-sea missions. In this paper, we report results and experiences from a messaging system that converts speech to text in a submarine, sends text messages to the surface, and reconstructs those messages as synthetic lip-synchronous videos of the speakers. The resulting system was tested during an actual dive to Titanic in the summer of 2022. We achieved an acceptable latency for a system of such complexity as well as good quality. The system demonstration video can be found at the following link: https://youtu.be/C4lyM86-5Ig
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。