通过屏蔽注意力头实现无需指令的音频任务可靠指定
AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
- 仅屏蔽解码器中部分注意力头,即可触发特定音频任务功能
- 在单任务和复合任务上表现优于或相当於使用指令的方法
- 揭示了大音频语言模型存在可被激活的特定功能路径
当前大型音频语言模型(LALMs)虽扩展了文本大语言模型(LLMs)的通用声学理解能力,但普遍存在提示敏感问题:相同意图的不同指令会导致结果差异显著。本文提出AHAMask方法,仅在LALMs的解码器仅有的LLM主干中屏蔽部分注意力头,即可在无指令情况下触发特定声学任务功能。该掩码通过在LALM上训练获得,可训练参数数量等于其LLM主干中的注意力头数。实验表明,采用这种选择性注意力头掩码,在单任务与复合任务上均达到或超越使用指令的效果。该方法不仅实现了对LALMs的可靠声学任务指定,还揭示了其注意力头中存在可被激活的‘功能通路’。
原文摘要 · Abstract (English)
Although current large audio language models (LALMs) extend text large language models (LLMs) with generic acoustic understanding abilities, they usually suffer from prompt sensitivity, where different instructions of the same intention can yield drastically different outcomes. In this work, we propose AHAMask, where we simply mask some of the attention heads in the decoder-only LLM backbone of LALMs, to trigger specific acoustic task functionalities without instructions. These masks are efficiently obtained by training on an LALM, with the number of trainable parameters equal to the attention head count in its LLM backbone. We show by experiments that applying such selective attention head masks achieves comparable or even better performance than using instructions, either on single or composite tasks. Besides achieving reliable acoustic task specification for LALMs, this also reveals that LALMs exhibit certain "functional pathways" in their attention heads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。