arXiv:2503.00356cs.CLcs.AI2025-03中稿 · Oral Presentation …

用预训练模型提升越南语事实验证准确率

BERT-based model for Vietnamese Fact Verification Dataset

  • 融合句子选择与分类的统一网络架构
  • 严格准确率达75.11%,比基线提升28.83%
  • 适用于越南语信息可信度检测场景

信息技术快速发展使信息获取更便捷,但也对信息真实性验证提出更高要求,尤其在越南语环境中。本文提出一种基于BERT的统一框架,整合句子选择与分类模块,利用预训练模型PhoBERT和XLM-RoBERTa作为主干网络,在越南语事实验证数据集ISE-DSC01上进行训练。实验表明,该模型在所有三个评估指标上均优于基线模型,严格准确率达到75.11%,较基线提升28.83%。

原文摘要 · Abstract (English)

The rapid advancement of information and communication technology has facilitated easier access to information. However, this progress has also necessitated more stringent verification measures to ensure the accuracy of information, particularly within the context of Vietnam. This paper introduces an approach to address the challenges of Fact Verification using the Vietnamese dataset by integrating both sentence selection and classification modules into a unified network architecture. The proposed approach leverages the power of large language models by utilizing pre-trained PhoBERT and XLM-RoBERTa as the backbone of the network. The proposed model was trained on a Vietnamese dataset, named ISE-DSC01, and demonstrated superior performance compared to the baseline model across all three metrics. Notably, we achieved a Strict Accuracy level of 75.11\%, indicating a remarkable 28.83\% improvement over the baseline model.

事实验证越南语BERT自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。