用AI把医学论文简化成初中生能懂的语言,效果比拼谁更简单准确。
UM_FHS at TREC 2024 PLABA: Exploration of Fine-tuning and AI agent approach for plain language adaptations of biomedical text
- 用提示工程和双AI代理策略生成通俗版本
- gpt-4o-mini在简洁性和整体表现上更优
- 适合需要快速理解医学内容的非专业人士
本文介绍我们在TREC 2024 PLABA赛道的提交工作,目标是将生物医学摘要简化为适合8年级学生(13-14岁)理解的通俗语言。我们测试了三种方法:基于OpenAI gpt-4o和gpt-4o-mini的提示工程基准、双AI代理方案以及微调模型。评估采用定性指标(5分李克特量表,涵盖简洁性、准确性、完整性与简明度)和定量可读性指标(Flesch-Kincaid年级水平、SMOG指数)。结果表明,使用gpt-4o-mini的提示工程方法在定性表现上优于双代理方案和gpt-4o微调模型;微调模型在准确性和完整性上更佳,但简洁性较差。最终发现,gpt-4o-mini的提示工程优于迭代改进的双代理策略及gpt-4o微调方案。后续将深入分析结果并探索更高级评估方法。
原文摘要 · Abstract (English)
This paper describes our submissions to the TREC 2024 PLABA track with the aim to simplify biomedical abstracts for a K8-level audience (13-14 years old students). We tested three approaches using OpenAI's gpt-4o and gpt-4o-mini models: baseline prompt engineering, a two-AI agent approach, and fine-tuning. Adaptations were evaluated using qualitative metrics (5-point Likert scales for simplicity, accuracy, completeness, and brevity) and quantitative readability scores (Flesch-Kincaid grade level, SMOG Index). Results indicated that the two-agent approach and baseline prompt engineering with gpt-4o-mini models show superior qualitative performance, while fine-tuned models excelled in accuracy and completeness but were less simple. The evaluation results demonstrated that prompt engineering with gpt-4o-mini outperforms iterative improvement strategies via two-agent approach as well as fine-tuning with gpt-4o. We intend to expand our investigation of the results and explore advanced evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。