arXiv:2605.06937cs.LG2026-05被引 1

让LLM做文献筛选更可靠:可复现的提示优化流程

A Reproducible Optimisation Protocol for Calibrating Prompt-Based Large Language Model Workflows in Evidence Synthesis

论文配图:A Reproducible Optimisation Protocol for Calibrating Prompt-Based Large Language Model Workflows in Evidence Synthesis
图 1 · 摘自论文原文
  • 分离任务规则与提示模板,用指标驱动优化
  • 小模型执行任务,大模型指导提示调优,提升准确率
  • 输出可检查的完整工作流,适合医疗文献综述场景

本文提出一种可复现的提示优化流程,用于结构化证据合成任务中的基于提示的大语言模型(LLM)。该方法将定义科学任务的规则与可变的提示框架分离,基于标注数据或参考示例及明确的任务指标对提示框架进行优化,并将校准后的工作流作为包含规范、指标、设置和评估轨迹的可检查产物保存。示例代码使用DSPy和GEPA工具实现该协议,但其核心逻辑可推广至支持结构化任务定义、指标引导搜索和产物复用的其他提示优化框架。以标题与摘要筛查为验证案例,因其具备标注基准数据和清晰评估指标。所演示的工作流采用小型学生模型执行科学任务,大型反思型模型在校准过程中引导提示优化。本文展示了编译、产物往返传递过程,以及优化预算对小型学生模型性能的影响。

原文摘要 · Abstract (English)

This methods article presents a reproducible calibration workflow for prompt-based large language models (LLMs) in structured evidence-synthesis tasks. The method separates the rules that define the scientific task from the mutable prompt harness that frames and applies them. It optimises that harness against labelled or reference examples and an explicit task metric, then preserves the calibrated workflow as an inspectable artefact with its specification, metric, settings, and evaluation traces. The example code instantiates the protocol with DSPy and GEPA tools, but the underlying logic can transfer to other prompt-optimisation frameworks that support structured task definitions, metric-guided search, and artefact reuse. Title and abstract screening is the worked validation case because it provides labelled benchmark data and clear evaluation metrics. The demonstrated workflow uses a smaller student LLM for performing the scientific task execution and a larger reflection LLM to steer the prompt optimisation process during calibration. This work shows compilation, artefact round-tripping, and how optimisation budget affects a smaller student model.

提示工程证据合成LLM优化可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。