让音频大模型学会识别知识边界,主动拒绝不懂的问题
Towards Reliable Large Audio Language Model
- 用多模态思维链和有监督微调提升模型可靠性
- 提出新评估指标RGI,发现两类方法均有效
- 可靠性可跨语音、音乐、通用声音迁移
近年来,大音频语言模型(LALMs)在语音、音乐和通用声音的通用理解与推理方面展现出令人瞩目的成果和前景。然而,这些模型仍缺乏识别自身知识边界并主动拒绝回答未知问题的能力。尽管已有研究提升了大语言模型的可靠性,但可靠的音频大模型仍鲜有探索。本文系统研究了多种实现可靠LALMs的方法,包括无需训练的多模态思维链(MCoT)以及基于训练的有监督微调(SFT)。此外,我们指出以往评估指标的局限性,提出新的可靠性增益指数(RGI)来衡量不同方法的有效性。结果表明,训练无关和训练相关的策略均能在不同程度上提升模型可靠性。更重要的是,我们发现可靠性是一种‘元能力’,可在语音、音乐和通用声音之间迁移,尽管三者在结构和内容上存在显著差异。
原文摘要 · Abstract (English)
Recent advancements in large audio language models (LALMs) have demonstrated impressive results and promising prospects in universal understanding and reasoning across speech, music, and general sound. However, these models still lack the ability to recognize their knowledge boundaries and refuse to answer questions they don't know proactively. While there have been successful attempts to enhance the reliability of LLMs, reliable LALMs remain largely unexplored. In this paper, we systematically investigate various approaches towards reliable LALMs, including training-free methods such as multi-modal chain-of-thought (MCoT), and training-based methods such as supervised fine-tuning (SFT). Besides, we identify the limitations of previous evaluation metrics and propose a new metric, the Reliability Gain Index (RGI), to assess the effectiveness of different reliable methods. Our findings suggest that both training-free and training-based methods enhance the reliability of LALMs to different extents. Moreover, we find that awareness of reliability is a "meta ability", which can be transferred across different audio modalities, although significant structural and content differences exist among sound, music, and speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。