构建新数据集并用图模型增强长音频会议理解
Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding

- 将音频特征建模为多维图,实现信息融合
- 提出代理规划机制,提升长期上下文记忆能力
- 专为长音频会议问答设计,适合语音理解研究者
长时音频会议理解(LAMU)日益受到关注,但特定任务的问答数据集仍稀缺。现有语音问答范式与先进语音大模型存在声学信息丢失和长程上下文记忆差的问题。为此,我们构建了LongAudioQA数据集,并提出GRGA模型,将异构音频特征建模为多维图,利用代理规划实现检索与答案生成,有效缓解信息丢失与记忆衰减问题。
原文摘要 · Abstract (English)
While long-form audio meeting understanding (LAMU) is garnering growing attention, task-specific question answering (QA) datasets remain scarce. Existing speech QA paradigms and state-of-the-art Speech LLMs suffer from acoustic information loss and poor long-term context memory. To address these issues, we construct the LongAudioQA dataset and propose the GRGA model, which models heterogeneous audio features into a multi-dimensional graph and leverages agent planning for retrieval and answer generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。