针对生物医学问答,按问题类型定制不同推理策略,提升答案准确性和证据可靠性。
From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b

- 按问题类型选用不同推理方法:是/否、事实型、列表型分别对应不同处理流程。
- 在BioASQ 14b任务中,事实型问题子任务获批次4第一名,整体表现具竞争力。
- 多智能体协作与自我反思机制,适合需要高可信度答案的医学研究场景。
生物医学问答不仅需从文献中准确提取信息,还需跨多篇文档可靠整合证据。本研究提出一种面向BioASQ 14b任务B的问题类型感知大语言模型框架,通过根据问题类型选择不同推理流程,提升答案鲁棒性与证据可追溯性。对于是/否类问题,采用片段打乱与自我反思以降低对证据顺序的敏感性;对于事实型问题,结合全片段输入与基于思维链的上下文学习,实现精准生物医学实体识别;对于列表型问题,引入多智能体架构,分阶段完成证据抽取、候选生成、答案验证与最终聚合。基于BioASQ 13b的预实验筛选有效策略后,该框架在正式的BioASQ 14b任务中表现优异,在多个批次中保持竞争力,并在批次4的事实型子任务中获得第一名。结果表明,结合问题类型感知推理、集成预测与智能体验证,能有效支持可靠的生物医学问答。
原文摘要 · Abstract (English)
Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration of evidence across multiple documents. This study presents a question-type-specific large language model (LLM) framework for BioASQ 14b Task B, designed to improve answer robustness and evidence grounding in biomedical question answering. Rather than applying a single prompting strategy to all questions, the framework selects different inference procedures for yes/no, factoid, and list questions according to their distinct reasoning and evaluation requirements. For yes/no questions, snippet shuffling and self-reflection are used to reduce sensitivity to evidence ordering and improve decision stability. For factoid questions, full-snippet input is combined with chain-of-thought-based in-context learning to support accurate biomedical entity identification. For list questions, a multi-agent architecture is employed, in which evidence extraction, candidate generation, answer verification, and final aggregation are handled collaboratively. Preliminary experiments on BioASQ 13b were used to identify effective inference strategies for each question type, and the resulting framework was subsequently evaluated in the official BioASQ 14b Task B challenge. In the official evaluation, our framework showed competitive performance across multiple batches and achieved first place in the factoid subtask of Batch 4. These results demonstrate the effectiveness of combining question-type-specific inference, ensemble prediction, and agent-based verification for reliable biomedical question answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。