揭秘多模态同步机制,为统一模型跨模态对齐提供新思路。
Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
- 用统一解码器架构研究视频、文本与语音的时序对齐机制。
- 视频优先提升本域性能,文本优先增强跨域泛化能力。
- 提出音素级评估指标,精准捕捉传统方法忽略的对齐误差。
统一的解码器仅变压器在多模态生成中展现出潜力,但其在异构采样率下的模态同步机制仍不明确。本文通过视频-文本-语音合成(VTTS)这一受控任务,研究了细粒度时序对齐问题。采用基于VoxCeleb2数据集训练的统一解码器模型Visatronic,探究:(i)各模态如何互补贡献;(ii)位置编码策略如何实现异构速率间的同步;(iii)模态顺序如何影响本域性能与跨域迁移的权衡;(iv)音素级同步指标如何诊断逐音素时序误差。结果表明,“全局序列索引”和“共时序索引”均能实现强同步性能,其中共时序索引无需显式时间戳即可达成对齐。文本保障可懂性,视频提供时间线索与情感表达。模态顺序呈现稳定权衡:视频优先提升本域表现,文本优先更利于跨域泛化。大规模训练揭示可迁移的同步策略。此外,引入音素级指标TimeSync,可发现帧级指标忽略的时序错位。这些发现确立了VTTS作为理解统一多模态解码器时序同步机制的重要测试平台。
原文摘要 · Abstract (English)
Unified decoder-only transformers have shown promise for multimodal generation, yet the mechanisms by which they synchronize modalities with heterogeneous sampling rates remain underexplored. We investigate these mechanisms through video-text-to-speech (VTTS) synthesis-a controlled task requiring fine-grained temporal alignment between sparse text, video, and continuous speech. Using a unified decoder-only transformer, dubbed Visatronic, trained on VoxCeleb2, we study: (i) how modalities contribute complementary information, (ii) how positional encoding strategies enable synchronization across heterogeneous rates, (iii) how modality ordering shapes the trade-off between in-domain performance and cross-domain transfer, (iv) how phoneme-level synchronization metrics provide diagnostic insight into per-phoneme timing errors. Our findings reveal that both "global sequential indexing'' (unique position IDs across modalities) and "co-temporal ordered indexing'' (identical IDs for temporally corresponding tokens) achieve strong synchronization performance, with co-temporal ordered indexing providing a simple mechanism without explicit timestamp metadata. Both text and video contribute complementary signals: text ensures intelligibility while video provides temporal cues and emotional expressiveness. Modality ordering reveals a consistent trade-off: video-first ordering achieves stronger in-domain performance while text-first ordering generalizes more robustly to unseen domains. Our findings also reveal, that diverse large-scale training enables transferable synchronization strategies. To enable fine-grained analysis, we also introduce TimeSync, a phoneme-level metric that reveals temporal misalignments overlooked by frame-level metrics. These insights establish VTTS as a valuable testbed for understanding temporal synchronization in unified multimodal decoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。