arXiv:2502.16778cs.LGcs.AI2025-02被引 2

研究生态网络结构特征在数据缺失下的稳定性,帮学者选更可靠的分析指标。

The Robustness of Structural Features in Species Interaction Networks

  • 用148个真实互作网络测试6种拓扑指标对缺边的敏感度。
  • 部分算法(如Louvain)社区检测结果变化更平缓,更抗数据缺失。
  • 不同互作类型下指标稳健性不同,为实际分析提供选型依据。

物种互作网络是描述生态群落的重要工具,通常以节点代表物种、边代表互作关系。为对相似网络群体进行抽象推断,生态学家常使用图拓扑指标总结结构特征。然而,获取底层数据困难,可能导致部分互作被遗漏,因此需了解不同结构指标受数据缺失的影响程度。本研究分析了148个现实世界的二分网络,涵盖四类互作关系(传粉、宿主-寄生、植物-蚂蚁、种子传播)。对每个网络,测量六种拓扑属性:连通分支数、节点介数方差、节点PageRank方差、最大特征值、非零特征值数量,以及四种算法(Clauset-Newman-Moore、Louvain、标签传播、Girvan-Newman)的社区检测结果。通过逐步添加模拟遗漏的边,评估这些属性的变化。结果发现,不同指标对数据缺失的鲁棒性差异显著;例如,Clauset-Newman-Moore和Louvain算法的社区检测随边增加呈现更平滑变化,表明其更稳健。此外,某些指标的鲁棒性还因互作类型而异。该研究为处理不完整生态网络数据时选择合适指标提供了依据。

原文摘要 · Abstract (English)

Species interaction networks are a powerful tool for describing ecological communities; they typically contain nodes representing species, and edges representing interactions between those species. For the purposes of drawing abstract inferences about groups of similar networks, ecologists often use graph topology metrics to summarize structural features. However, gathering the data that underlies these networks is challenging, which can lead to some interactions being missed. Thus, it is important to understand how much different structural metrics are affected by missing data. To address this question, we analyzed a database of 148 real-world bipartite networks representing four different types of species interactions (pollination, host-parasite, plant-ant, and seed-dispersal). For each network, we measured six different topological properties: number of connected components, variance in node betweenness, variance in node PageRank, largest Eigenvalue, the number of non-zero Eigenvalues, and community detection as determined by four different algorithms. We then tested how these properties change as additional edges -- representing data that may have been missed -- are added to the networks. We found substantial variation in how robust different properties were to the missing data. For example, the Clauset-Newman-Moore and Louvain community detection algorithms showed much more gradual change as edges were added than the label propagation and Girvan-Newman algorithms did, suggesting that the former are more robust. Robustness also varied for some metrics based on interaction type. These results provide a foundation for selecting network properties to use when analyzing messy ecological network data.

生态网络拓扑分析数据缺失社区检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。