用大模型生成问题并检索证据,自动判断新闻可信度。
From Questions to Trust Reports: A LLM-IR Framework for the TREC 2025 DRAGUN Track
- 结合大模型提问与语义筛选,提升证据召回相关性。
- 基于思维链扩展查询,显著提升相关性和可信度评分。
- 适合关注虚假信息检测与可信度评估的研究者。
TREC 2025 DRAGUN赛道聚焦于提升用户对网络新闻可信度的评估能力。我们提交的UR_Trecking系统参与任务1(关键问题生成)和任务2(检索增强型可信度报告生成)。方法融合大模型问题生成、语义过滤、聚类多样性控制及多种查询扩展策略(包括基于思维链的Chain-of-Thought扩展),从MS MARCO V2.1分段语料库中检索相关证据。通过monoT5模型重排序,并结合大模型相关性判断与领域可信度数据集进行过滤。任务2中,选定证据由大模型合成带引用的简洁可信度报告。官方评测结果显示,思维链查询扩展和重排序显著提升相关性与领域可信度表现,而问题生成质量中等,仍有改进空间。最后总结了主要挑战,并提出未来提升系统鲁棒性与可信度评估能力的方向。
原文摘要 · Abstract (English)
The DRAGUN Track at TREC 2025 targets the growing need for effective support tools that help users evaluate the trustworthiness of online news. We describe the UR_Trecking system submitted for both Task 1 (critical question generation) and Task 2 (retrieval-augmented trustworthiness reporting). Our approach combines LLM-based question generation with semantic filtering, diversity enforcement using clustering, and several query expansion strategies (including reasoning-based Chain-of-Thought expansion) to retrieve relevant evidence from the MS MARCO V2.1 segmented corpus. Retrieved documents are re-ranked using a monoT5 model and filtered using an LLM relevance judge together with a domain-level trustworthiness dataset. For Task 2, selected evidence is synthesized by an LLM into concise trustworthiness reports with citations. Results from the official evaluation indicate that Chain-of-Thought query expansion and re-ranking substantially improve both relevance and domain trust compared to baseline retrieval, while question-generation performance shows moderate quality with room for improvement. We conclude by outlining key challenges encountered and suggesting directions for enhancing robustness and trustworthiness assessment in future iterations of the system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。