构建可复用的自动化评估工具,助力新闻可信度辅助系统评测。
Resources for Automated Evaluation of Assistive RAG Systems that Help Readers with News Trustworthiness Assessment
- 开发自动评分系统,基于人类标注的评判标准评估报告质量。
- 在250词报告任务中,自动化评分与人工评分相关性达0.872。
- 适合研究新闻可信度评估、RAG系统评测的学者使用。
如今许多读者难以判断网络新闻的可信度,因真实报道与虚假信息并存。TREC 2025 DRAGUN(新闻理解中的检测、检索与增强生成)赛道为研究人员提供了开发和评估辅助性RAG系统的机会,以生成面向读者、有据可依的新闻可信度报告。该赛道包含两项任务:任务1为问题生成,生成10个排序的调查性问题;任务2为主要任务,基于MS MARCO V2.1 Segmented Corpus生成250词的报告。我们邀请TREC评估员针对30篇新闻文章制定重要性加权的问答评判标准,代表评估员认为读者判断可信度所需的关键信息。评估员随后依据这些标准对参赛团队提交的结果进行人工打分。为使任务与评判标准可复用,我们建立了自动化评估流程,可对非原参赛结果进行评分。实验表明,该自动评分系统在任务1和任务2上的评估结果与人工评估的相关性分别为Kendall's τ=0.678和τ=0.872。这些资源不仅支持对辅助新闻可信度评估的RAG系统的评估,还为提升自动化评估方法的研究提供了基准。
原文摘要 · Abstract (English)
Many readers today struggle to assess the trustworthiness of online news because reliable reporting coexists with misinformation. The TREC 2025 DRAGUN (Detection, Retrieval, and Augmented Generation for Understanding News) Track provided a venue for researchers to develop and evaluate assistive RAG systems that support readers' news trustworthiness assessment by producing reader-oriented, well-attributed reports. As the organizers of the DRAGUN track, we describe the resources that we have newly developed to allow for the reuse of the track's tasks. The track had two tasks: (Task 1) Question Generation, producing 10 ranked investigative questions; and (Task 2, the main task) Report Generation, producing a 250-word report grounded in the MS MARCO V2.1 Segmented Corpus. As part of the track's evaluation, we had TREC assessors create importance-weighted rubrics of questions with expected short answers for 30 different news articles. These rubrics represent the information that assessors believe is important for readers to assess an article's trustworthiness. The assessors then used their rubrics to manually judge the participating teams' submitted runs. To make these tasks and their rubrics reusable, we have created an automated process to judge runs not part of the original assessing. We show that our AutoJudge ranks existing runs well compared to the TREC human-assessed evaluation (Kendall's $τ= 0.678$ for Task 1 and $τ= 0.872$ for Task 2). These resources enable both the evaluation of RAG systems for assistive news trustworthiness assessment and, with the human evaluation as a benchmark, research on improving automated RAG evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。