arXiv:2409.01201eess.AScs.AI2024-09中稿 · DCASE2024 Workshop被引 7

优化音频字幕模型EnCLAP,提升生成质量。

EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance

  • 改进声学编码器结构,增强音频特征提取能力。
  • 在更大规模数据集上预训练,显著提升字幕准确率。
  • 引入重排序机制,改善生成结果的语义连贯性。

本文旨在分析并优化最先进的自动化音频字幕模型EnCLAP。通过系统研究声学编码器组件的修改、不同规模数据集上的预训练效果,以及重排序方案的有效性,我们进行了大量实验与生成字幕的定量分析。基于这些发现,提出EnCLAP++,其性能显著优于原始模型,在多个评估指标上实现明显提升。该工作为音频字幕生成提供了可复现的优化路径。

原文摘要 · Abstract (English)

In this work, we aim to analyze and optimize the EnCLAP framework, a state-of-the-art model in automated audio captioning. We investigate the impact of modifying the acoustic encoder components, explore pretraining with different dataset scales, and study the effectiveness of a reranking scheme. Through extensive experimentation and quantitative analysis of generated captions, we develop EnCLAP++, an enhanced version that significantly surpasses the original.

音频字幕模型优化深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。