arXiv:2505.07440cs.CL2025-05

用弱监督方法为24个产业匹配常见任务,补全常识知识库。

Matching Tasks with Industry Groups for Augmenting Commonsense Knowledge

  • 训练神经模型学习任务与产业的亲和度,通过聚类筛选每产业前k个任务。
  • 从新闻数据中提取2339条产业-任务三元组,准确率达0.86。
  • 结果可直接补充现有常识知识库,适合知识图谱构建者使用。

常识知识库(KB)是广泛用于提升机器学习应用的专业知识来源。然而,即使对于大型知识库如ConceptNet,捕捉各产业领域的显式知识仍具挑战性。例如,ConceptNet中关于各类产业执行的通用任务样本极少。本文旨在填补这一空白,提出一种弱监督框架,用于增强常识知识库中产业群体(IG)所执行的任务信息。我们尝试通过训练神经模型学习任务与产业的亲和度,并结合聚类方法,为每个产业选取前k个最相关的任务。基于两个公开新闻数据集,我们为24个产业群体提取了总计2339条形式为⟨IG, is capable of, task⟩的三元组,准确率达到0.86。该结果验证了所提取任务-产业对的可靠性,可直接添加至现有知识库中。

原文摘要 · Abstract (English)

Commonsense knowledge bases (KB) are a source of specialized knowledge that is widely used to improve machine learning applications. However, even for a large KB such as ConceptNet, capturing explicit knowledge from each industry domain is challenging. For example, only a few samples of general {\em tasks} performed by various industries are available in ConceptNet. Here, a task is a well-defined knowledge-based volitional action to achieve a particular goal. In this paper, we aim to fill this gap and present a weakly-supervised framework to augment commonsense KB with tasks carried out by various industry groups (IG). We attempt to {\em match} each task with one or more suitable IGs by training a neural model to learn task-IG affinity and apply clustering to select the top-k tasks per IG. We extract a total of 2339 triples of the form $\langle IG, is~capable~of, task \rangle$ from two publicly available news datasets for 24 IGs with the precision of 0.86. This validates the reliability of the extracted task-IG pairs that can be directly added to existing KBs.

常识知识产业知识弱监督知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。