用可信知识库提升未知信息过滤准确率
TrustDataFilter:Leveraging Trusted Knowledge Base Data for More Effective Filtering of Unknown Information
- 通过对比可信知识与待过滤数据,推理其逻辑关系
- 在生物、辐射、科学领域实验中显著提升过滤一致性
- 适配大模型插件式使用,适合知识库构建场景
随着技术进步和市场需求变化,领域专用知识库的构建日益重要,常依赖网络爬虫或行业数据库,导致数据准确性与一致性问题。针对领域数据内部关联性强的特点,提出Self-NLI-TDF框架,通过将可信过滤知识与待过滤数据进行自然语言推理比较,推断二者间逻辑关系,从而提升过滤性能。该框架采用可插拔的大语言模型进行可信度评估,并使用NLI领域预训练的RoBERTa-MNLI模型进行推理。在生物学、辐射、科学三个领域构建了三个数据集,分别使用RoBERTa、GPT3.5和本地Qwen2模型进行实验,结果表明该框架有效提升了过滤质量,生成结果更具一致性和可靠性。
原文摘要 · Abstract (English)
With the advancement of technology and changes in the market, the demand for the construction of domain-specific knowledge bases has been increasing, either to improve model performance or to promote enterprise innovation and competitiveness. The construction of domain-specific knowledge bases typically relies on web crawlers or existing industry databases, leading to problems with accuracy and consistency of the data. To address these challenges, we considered the characteristics of domain data, where internal knowledge is interconnected, and proposed the Self-Natural Language Inference Data Filtering (self-nli-TDF) framework. This framework compares trusted filtered knowledge with the data to be filtered, deducing the reasoning relationship between them, thus improving filtering performance. The framework uses plug-and-play large language models for trustworthiness assessment and employs the RoBERTa-MNLI model from the NLI domain for reasoning. We constructed three datasets in the domains of biology, radiation, and science, and conducted experiments using RoBERTa, GPT3.5, and the local Qwen2 model. The experimental results show that this framework improves filter quality, producing more consistent and reliable filtering results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。