用多智能体融合方法提升医疗问答论坛的多视角摘要质量
YaleNLP @ PerAnsSumm 2025: Multi-Perspective Integration via Mixture-of-Agents for Enhanced Healthcare QA Summarization
- 设计多层反馈聚合的混合智能体框架,整合多个大模型输出
- 在多视角识别任务上,模型表现比基线提升28%至0.51
- 适合需要综合多方观点的医疗信息摘要场景
自动化生成医疗社区问答论坛的摘要面临挑战,因每个问题下存在多种用户观点。为此,PerAnsSumm共享任务旨在从不同回答中识别视角并生成综合答案。本研究采用两种互补范式:一是基于QLoRA微调LLaMA-3.3-70B-Instruct的训练方法;二是零样本与少样本提示(使用LLaMA-3.3-70B-Instruct和GPT-4o)及混合智能体(MoA)框架,通过多层反馈聚合整合多样大模型输出。在视角跨度识别任务中,GPT-4o零样本得分0.57,显著优于基线模型的0.40;2层MoA配置使LLaMA性能提升28%至0.51。在基于视角的摘要任务中,GPT-4o零样本得分为0.42,优于最佳的LLaMA零样本结果(0.28),2层MoA将LLaMA性能提升32%至0.37。此外,在少样本设置下,基于sentence-transformer的示例选择比人工选取更有效,但对GPT-4o的提示增益有限。YaleNLP团队在共享任务中获得总排名第二。
原文摘要 · Abstract (English)
Automated summarization of healthcare community question-answering forums is challenging due to diverse perspectives presented across multiple user responses to each question. The PerAnsSumm Shared Task was therefore proposed to tackle this challenge by identifying perspectives from different answers and then generating a comprehensive answer to the question. In this study, we address the PerAnsSumm Shared Task using two complementary paradigms: (i) a training-based approach through QLoRA fine-tuning of LLaMA-3.3-70B-Instruct, and (ii) agentic approaches including zero- and few-shot prompting with frontier LLMs (LLaMA-3.3-70B-Instruct and GPT-4o) and a Mixture-of-Agents (MoA) framework that leverages a diverse set of LLMs by combining outputs from multi-layer feedback aggregation. For perspective span identification/classification, GPT-4o zero-shot achieves an overall score of 0.57, substantially outperforming the 0.40 score of the LLaMA baseline. With a 2-layer MoA configuration, we were able to improve LLaMA performance up by 28 percent to 0.51. For perspective-based summarization, GPT-4o zero-shot attains an overall score of 0.42 compared to 0.28 for the best LLaMA zero-shot, and our 2-layer MoA approach boosts LLaMA performance by 32 percent to 0.37. Furthermore, in few-shot setting, our results show that the sentence-transformer embedding-based exemplar selection provides more gain than manually selected exemplars on LLaMA models, although the few-shot prompting is not always helpful for GPT-4o. The YaleNLP team's approach ranked the overall second place in the shared task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。