分析新闻中事实如何通过引用来源增强可信度
FactAppeal: Identifying Epistemic Factual Appeals in News Media
- 标注3226条新闻语句,识别事实陈述与支撑来源的关联
- 模型在90亿参数下达到0.73的宏平均F1分数
- 适合研究事实核查、媒体可信度与自然语言推理的人
事实如何获得可信性?我们提出新的任务——认识论引用识别,即判断并解析事实陈述是否依赖外部来源或证据。为推动该任务研究,我们构建了FactAppeal数据集,包含3,226条英文新闻句子的手动标注。不同于以往仅关注事实检测与验证的资源,FactAppeal揭示了事实背后的精细认识论结构与证据基础。标注包含细粒度特征:事实陈述范围、来源类型(如当事人、目击者、专家、直接证据)、是否具名、角色与认知资质说明、引用方式(直接或间接引述)等。我们采用20亿至90亿参数范围的编码器与生成解码器模型进行建模。最佳模型基于Gemma 2 9B,在测试集上取得0.73的宏平均F1得分。
原文摘要 · Abstract (English)
How is a factual claim made credible? We propose the novel task of Epistemic Appeal Identification, which identifies whether and how factual statements have been anchored by external sources or evidence. To advance research on this task, we present FactAppeal, a manually annotated dataset of 3,226 English-language news sentences. Unlike prior resources that focus solely on claim detection and verification, FactAppeal identifies the nuanced epistemic structures and evidentiary basis underlying these claims and used to support them. FactAppeal contains span-level annotations which identify factual statements and mentions of sources on which they rely. Moreover, the annotations include fine-grained characteristics of factual appeals such as the type of source (e.g. Active Participant, Witness, Expert, Direct Evidence), whether it is mentioned by name, mentions of the source's role and epistemic credentials, attribution to the source via direct or indirect quotation, and other features. We model the task with a range of encoder models and generative decoder models in the 2B-9B parameter range. Our best performing model, based on Gemma 2 9B, achieves a macro-F1 score of 0.73.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。