arXiv:2601.04197cs.CL2026-01

无监督构建中文动词搭配数据库,提升语言模型的可解释性。

Automatic Construction of Chinese Verb Collostruction Database

  • 将动词搭配定义为有向无环图,用聚类算法从语料中自动提取
  • 生成的搭配具备功能独立性和典型性分级特征,统计验证有效
  • 基于搭配的最大匹配纠错法优于大模型,适合需解释性的场景

本文提出一种完全无监督的方法,用于构建中文动词搭配数据库,旨在通过提供明确且可解释的规则,弥补大语言模型在需要解释性和可解释性场景中的不足。论文将动词搭配形式化定义为投影、有根、有序且无环的有向图,并利用一系列聚类算法,从大规模语料库中检索出的句子列表中为给定动词生成搭配结构。统计分析表明,生成的搭配具有功能独立性和等级典型性特征。在动词语法错误纠正任务上的评估显示,基于最大匹配与搭配的纠错算法性能优于大语言模型。

原文摘要 · Abstract (English)

This paper proposes a fully unsupervised approach to the construction of verb collostruction database for Chinese language, aimed at complementing LLMs by providing explicit and interpretable rules for application scenarios where explanation and interpretability are indispensable. The paper formally defines a verb collostruction as a projective, rooted, ordered, and directed acyclic graph and employs a series of clustering algorithms to generate collostructions for a given verb from a list of sentences retrieved from large-scale corpus. Statistical analysis demonstrates that the generated collostructions possess the design features of functional independence and graded typicality. Evaluation with verb grammatical error correction shows that the error correction algorithm based on maximum matching with collostructions achieves better performance than LLMs.

自然语言处理搭配数据库可解释性无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。