让健康谣言纠错笔记会学习,越用越准。
Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes

- 用细粒度反馈记忆积累纠错经验,自动优化分析与写作
- 90%以上生成笔记优于人工,82%未评分帖也能有效纠错
- 从13小时缩短至2分钟,适合大规模健康信息治理
大型语言模型增强的社区笔记为社交媒体上及时、基于证据地纠正健康谣言提供了可扩展路径。然而,每次新帖子都重置,先前纠错经验无法复用。我们提出EvoNote,一种通过持续积累历史纠错案例的经验记忆实现自进化的能力框架。其核心是细粒度信用分配:将轨迹级反馈归因于健康相关笔记质量,并提炼为行动级记忆,用于论点分析、证据获取和笔记撰写。我们在包含1200个实例的多模态基准MM-HealthCN上评估,该数据集包含用户标记的健康帖子、人工撰写的社区笔记及众包的帮助性标签。在经人工验证的分层效用评判中,EvoNote生成的笔记在89.6%的情况下优于对应人工笔记;在无众包帮助性判断的“需更多评分”帖子中,仍能对82.0%的案例生成有效纠正内容。同时,候选纠正生成时间从人工流程的13小时以上降至2分钟以内。分析表明,性能提升源于更强的证据使用和可复用的纠错策略,证明自进化笔记生成是健康谣言治理的有前景范式。
原文摘要 · Abstract (English)
Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social platforms. However, they still reset at every post, leaving useful correction experience from prior cases unused. We introduce EvoNote, an agentic framework that enables health Community Notes generation to self-evolve through an evolving experience memory of prior misinformation correction episodes. Its core is fine-grained credit assignment: EvoNote grounds trajectory-level feedback in health-specific note qualities and distills it into action-level memory for claim analysis, evidence acquisition, and note writing. We evaluate EvoNote on MM-HealthCN, a 1.2K-instance multimodal benchmark of user-flagged health posts with human-written Community Notes and crowd-derived helpfulness labels. Under a human-validated hierarchical utility judge, EvoNote-generated notes are preferred over corresponding human-written notes in 89.6% of cases; on a separate set of Needs More Ratings posts without a crowd helpfulness verdict, EvoNote produces helpful notes for 82.0% of cases. It also reduces the median time needed to produce a candidate correction from over 13 hours in the human-note pipeline to under 2 minutes. Analyses link these gains to stronger evidence use and reusable correction strategies, positioning self-evolving note generation as a promising paradigm for health misinformation governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。