用GPT-4帮记者识别科学术语,还能自适应个性化解释。
De-jargonizing Science for Journalists with GPT-4: A Pilot Study
- 结合GPT-4与检索增强生成,自动识别科学摘要中的术语。
- 仅用摘要生成定义,准确率和质量反而高于使用全文上下文。
- 适合需要简化复杂文献的科技记者或科普工作者。
本研究评估了一种人机协同系统,利用GPT-4与检索增强生成(RAG)技术,基于读者自报知识水平识别并定义科学摘要中的专业术语。系统在识别术语方面表现出较高召回率,并保留了读者间对术语理解差异的相对关系,表明个性化辅助具有可行性。令人意外的是,仅以摘要为上下文生成定义,其准确性和质量略优于使用文章全文信息的RAG方法。结果表明生成式AI在辅助科学传播方面具有潜力,可为未来简化密集型文档工具的研发提供参考。
原文摘要 · Abstract (English)
This study offers an initial evaluation of a human-in-the-loop system leveraging GPT-4 (a large language model or LLM), and Retrieval-Augmented Generation (RAG) to identify and define jargon terms in scientific abstracts, based on readers' self-reported knowledge. The system achieves fairly high recall in identifying jargon and preserves relative differences in readers' jargon identification, suggesting personalization as a feasible use-case for LLMs to support sense-making of complex information. Surprisingly, using only abstracts for context to generate definitions yields slightly more accurate and higher quality definitions than using RAG-based context from the fulltext of an article. The findings highlight the potential of generative AI for assisting science reporters, and can inform future work on developing tools to simplify dense documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。