arXiv:2606.07548cs.IRcs.AI2026-06被引 1

精心设计提示词能让小模型实现大模型的推理效果

Evaluating Advanced Prompting on Gemini Flash for Multi-Hop Biomedical QA

  • 用角色扮演+分步思考+格式约束构建复杂提示
  • 最佳提示使得分从0.565提升至0.720
  • 高效模型在优秀提示下接近下一代性能

MedHopQA挑战对大型语言模型(LLMs)提出了严峻考验:在高风险生物医学领域进行复杂多跳推理。本文通过直接调用API评估谷歌Gemini Flash系列模型,重点研究先进提示工程的影响。我们为Gemini 2.0 Flash设计了一种多组件复杂提示,融合角色扮演、显式多示例思维链(CoT)及详细格式规则。最佳实验结果达到概念层级得分0.720,显著优于基线提示的0.565。令人惊讶的是,该高效模型在优秀提示下的表现几乎与下一代Gemini 2.5 Flash相当。研究证明,精巧的提示设计是激发现代LLM全部推理能力的关键因素。

原文摘要 · Abstract (English)

The MedHopQA challenge presents a critical test for Large Language Models (LLMs): complex, multi-hop reasoning in the high-stakes biomedical domain. This paper details our direct API-based evaluation of Google's Gemini Flash models, focusing on the impact of advanced prompt engineering. We designed a sophisticated, multi-component prompt for Gemini 2.0 Flash that combined role-playing, explicit multi-shot Chain-of-Thought (CoT) examples, and detailed formatting rules. Our best run, using this complex prompt, achieved a Concept Level Score of 0.720. This result dramatically outperformed a baseline prompt which scored only 0.565. Remarkably, this performance on the efficient Gemini 2.0 Flash was almost identical to the result from the next-generation Gemini 2.5 Flash. Our findings demonstrate that sophisticated prompt design is a critical factor for unlocking the full reasoning capabilities of modern LLMs.

提示工程生物医学多跳推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。