用迭代优化指南提升大模型标注质量,尤其在生物医学领域效果显著
Refining and Reusing Annotation Guidelines for LLM Annotation
- 通过迭代模拟标注流程,系统化重用与优化标注规范
- 在三个生物医学命名实体识别任务中准确率显著提升
- 适合需要高质量标注的医疗、科研类AI项目
尽管大语言模型在零样本标注任务中表现优异,但在标准金标数据集的特殊标注规范上仍存在困难。本文提出将标注指南进行系统性重用与迭代优化,构建一个模拟标注初期阶段的迭代审核框架。我们验证了三个假设:(1)指南整合的有效性,(2)推理优化模型的优势,(3)低监督下审核的可行性。在生物医学命名实体识别任务(NCBI Disease、BC5CDR、BioRED)上,对GPT、Gemini、DeepSeek三类大模型进行测试,结果证实了所有假设。虽然该框架在指南优化方面展现出良好潜力,但分析也揭示了仍有较大改进空间。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) demonstrate remarkable performance on zero-shot annotation tasks, they often struggle with the specialized conventions of gold-standard benchmarks. We propose the systematic reuse and refinement of annotation guidelines as an alignment mechanism, introducing an iterative moderation framework that simulates the early phases of annotation projects. We evaluate three hypotheses: (1) the efficacy of guideline integration, (2) the advantage of reasoning optimized models, and (3) the viability of moderation under minimal supervision. Testing across biomedical NER tasks (NCBI Disease, BC5CDR, BioRED) with three LLM families (GPT, Gemini, DeepSeek), our results empirically confirm all three hypotheses. While the iterative moderation framework shows good potential in effectively refining guidelines, our analysis also reveals substantial room for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。