arXiv:2607.06542cs.CL2026-07

无需标注数据,也能评估非人类序列的依赖解析可行性。

On the feasibility of dependency parsing of non-human sequences without a gold standard. Is evaluation possible in other species?

  • 利用网络科学分析序列长度分布,推断解析正确率下限。
  • 非人灵长类发声序列因长度衰减快,解析准确率可被可靠估计。
  • 该方法为动物交流研究提供无标注评估新路径,适合跨物种语言学研究者。

依赖解析旨在为序列寻找树状结构表示。无监督依赖解析试图在无真实标注的情况下训练解析模型。在人类语言中,可通过真实标注评估模型性能;但在其他物种中,真实标注未知,因此常认为无法评估无监督解析器的准确性,进而质疑非人类序列依赖解析的可行性。然而,本文借助网络科学最新进展,证明:由于非人灵长类发声或手势序列的长度分布衰减迅速,解析器所恢复的正确边比例必然较高。相比之下,人类语言序列不具备这一特性。因此,在非人灵长类中,无需真实标注即可实现有效评估,而在人类语言中则极为困难。

原文摘要 · Abstract (English)

Dependency parsing consists of finding a tree representation for a sequence. Unsupervised dependency parsing aims to develop parsing methods without a gold standard during model training. In human languages, an unsupervised parser can be evaluated because some gold standard is usually available or can be created. For other species, a gold standard is unknown. Thus one may conclude that it is impossible to determine the accuracy of an unsupervised parser and, consequently, dependency parsing is unfeasible in other species. However, here we apply recent advances in network science to demonstrate that the proportion of correct edges retrieved by a parser must be high for the sequences of vocalizations or gestures that non-human primates produce due to the fast decay of the sequence length distribution. In contrast, human language sequences lack that property. Therefore, evaluation without a gold standard is feasible in non-human primates but a hard problem in humans.

依赖解析无监督学习动物交流网络科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。