用对话续写训练大模型,让音频理解更懂指令。
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
- 让模型像接话一样续写对话,理解音频深层含义
- 零样本下可执行未见过的指令,无需多任务微调
- 在多个数据集上验证了对新指令的强泛化能力
我们提出一种基于大语言模型对话续写能力的指令跟随音频理解模型。该方法不直接生成训练数据中的目标描述,而是训练模型以输入描述引发对话的方式生成回复。这种对话续写训练有效缓解了描述多样性问题,能捕捉描述背后的真实语义。结果表明,仅在音频描述数据集上训练的模型,即可实现零样本指令跟随能力,无需多任务指令微调。在AudioCaps、WavCaps和Clotho数据集上,结合AudioBench音频场景问答测试,验证了模型对各类未见指令的有效响应能力。
原文摘要 · Abstract (English)
We propose an instruction-following audio comprehension model that leverages the dialogue continuation ability of large language models (LLMs). Instead of directly generating target captions in training data, the proposed method trains a model to produce responses as if the input caption triggered a dialogue. This dialogue continuation training mitigates the caption variation problem. Learning to continue a dialogue effectively captures the caption's meaning beyond its surface-level words. As a result, our model enables zero-shot instruction-following capability without multitask instruction tuning, even trained solely on audio captioning datasets. Experiments on AudioCaps, WavCaps, and Clotho datasets with AudioBench audio-scene question-answering tests demonstrate our model's ability to follow various unseen instructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。