arXiv:2602.16843cs.CL2026-02中稿 · 2nd LoResLM at EAC…被引 2

无需参考摘要,自动评估孟加拉语摘要的事实一致性。

BanglaSummEval: Reference-Free Factual Consistency Evaluation for Bangla Summarization

  • 基于问答自动生成评估问题,统一使用多语言指令微调模型处理全流程。
  • 在300份教育医疗领域摘要上与专家评分相关性达0.76(斯皮尔曼)。
  • 适合低资源语言摘要评估,提供可解释的诊断过程。

事实一致性评估对可靠文本摘要至关重要,尤其在医疗和新闻等高风险领域。然而,现有大多数评估指标忽视孟加拉语这一广泛使用但资源匮乏的语言,且常依赖参考摘要。我们提出 BanglaSummEval,一种无参考、基于问答的事实一致性评估框架。该方法通过从源文档和摘要中自动生成问题与答案,评估事实准确性和内容覆盖度。单一多语言指令微调语言模型完成问题生成、问答、候选答案提取及问题重要性加权。统一设计降低系统复杂度与计算成本。为捕捉表面重叠之外的语义一致性,采用 BERTScore-Recall 进行答案对比。在300份人工撰写的教育与医学领域摘要上验证,与专家评分呈现强相关性(皮尔逊相关系数 $r = 0.694$,斯皮尔曼等级相关系数 $ρ = 0.763$)。通过提供可解释的分步诊断与可靠评分,BanglaSummEval 为低资源语言环境下的事实一致性评估提供了实用且透明的解决方案。

原文摘要 · Abstract (English)

Evaluating factual consistency is essential for reliable text summarization, particularly in high-stakes domains such as healthcare and news. However, most existing evaluation metrics overlook Bangla, a widely spoken yet under-resourced language, and often depend on reference summaries. We introduce BanglaSummEval, a reference-free, question-answering-based framework for evaluating factual consistency in Bangla summarization. The proposed method assesses both factual accuracy and content coverage through automatically generated questions and answers derived from the source document and the summary. A single multilingual instruction-tuned language model handles question generation, question answering, candidate answer extraction, and question importance weighting. This unified design reduces system complexity and computational cost. To capture semantic consistency beyond surface-level overlap, we use BERTScore-Recall for answer comparison. We validate BanglaSummEval on 300 human-written summaries from educational and medical domains, demonstrating strong correlation with expert human judgments (Pearson's $r = 0.694$, Spearman's $ρ= 0.763$). By providing interpretable, step-wise diagnostics alongside reliable evaluation scores, BanglaSummEval offers a practical and transparent solution for factual consistency evaluation in low-resource language settings.

摘要评估事实一致性低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。