用自上而下的图算法自动构建跨语言语义地图,省时高效。
A Top-down Graph-based Tool for Modeling Classical Semantic Maps: A Crosslinguistic Case Study of Supplementary Adverbs
- 基于密集图与最大生成树,自上而下生成语义网络。
- 在副词跨语言研究中,效果接近人工标注,速度更快。
- 适合语言学研究者快速构建语义地图,支持可复现分析。
语义地图模型(SMMs)基于连接性假说,从跨语言实例中构建类网络的概念空间,广泛用于比较不同语言间的概念相似性和蕴含关系。然而,现有SMMs多依赖人工专家通过自下而上的方式构建,耗时费力。本文提出一种新型基于图的算法,以自上而下的方式自动生成概念空间与SMMs。算法首先构建稠密图,再依据我们提出的评估指标进行剪枝,形成最大生成树。这些指标结合了内在结构特征与外在表现,权衡精度与覆盖范围。在跨语言补充副词的案例研究中,该方法在效果和效率上均优于人工标注及其他自动化方法。工具已开源:https://github.com/RyanLiut/SemanticMapModel。
原文摘要 · Abstract (English)
Semantic map models (SMMs) construct a network-like conceptual space from cross-linguistic instances or forms, based on the connectivity hypothesis. This approach has been widely used to represent similarity and entailment relationships in cross-linguistic concept comparisons. However, most SMMs are manually built by human experts using bottom-up procedures, which are often labor-intensive and time-consuming. In this paper, we propose a novel graph-based algorithm that automatically generates conceptual spaces and SMMs in a top-down manner. The algorithm begins by creating a dense graph, which is subsequently pruned into maximum spanning trees, selected according to metrics we propose. These evaluation metrics include both intrinsic and extrinsic measures, considering factors such as network structure and the trade-off between precision and coverage. A case study on cross-linguistic supplementary adverbs demonstrates the effectiveness and efficiency of our model compared to human annotations and other automated methods. The tool is available at https://github.com/RyanLiut/SemanticMapModel.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。