arXiv:2603.23286cs.CV2026-03

通过拓扑结构评估绳结分类模型,揭示外观偏差的干扰并验证结构先验的有效性。

Physical Knot Classification Beyond Accuracy: A Benchmark and Diagnostic Study

  • 构建10类1440张图像数据集,用松结训练、紧结测试检验拓扑理解能力。
  • 拓扑距离可预测跨类别混淆,验证了拓扑感知评估框架的有效性。
  • 提出结构监督新方法,显著提升模型对真实结构的识别特异性,适合注重泛化性的研究者。

物理绳结分类是一项具有挑战性的细粒度识别任务,其关键判别线索是绳索交叉结构;然而,高闭集准确率可能源于低层外观捷径,而非真正的拓扑理解。本文引入一个包含1,440张图像、10个类别的数据集,以松散打结作为训练数据,紧密打结作为测试配置,探究结构引导训练是否带来拓扑特异性提升。结果表明,拓扑距离能有效预测多种主干网络下的残余类间混淆,验证了拓扑感知评估框架的实用性。进一步提出拓扑感知质心对齐(TACA)和辅助交叉数预测目标两种互补的结构监督方式。值得注意的是,使用TACA的Swin-T在标准协议下所有随机种子上均实现+1.18个百分点的特异性增益,而交叉数预测在不同数据规模下表现稳健,未出现质心对齐中的真实-随机反转现象。因果探针显示,背景变化即可导致17%-32%的预测翻转,手机照片准确率下降58%-69个百分点,凸显外观偏差仍是部署的主要障碍。这些结果共同表明,本诊断流程为评估手工设计的结构先验是否带来真实任务收益提供了系统且实用的工具。

原文摘要 · Abstract (English)

Physical knot classification is a challenging fine-grained recognition task in which the intended discriminative cue is rope crossing structure; however, high closed-set accuracy may still arise from low-level appearance shortcuts rather than genuine topological understanding. In this work, we introduce dataset (1,440 images, 10 classes), which trains models on loosely tied knots and evaluates them on tightly dressed configurations to probe whether structure-guided training yields topology-specific gains. We demonstrate that topological distance successfully predicts residual inter-class confusion across multiple backbone architectures, validating the utility of our topology-aware evaluation framework. Furthermore, we propose topology-aware centroid alignment (TACA) and an auxiliary crossing-number prediction objective as two complementary forms of structural supervision. Notably, Swin-T with TACA achieves a consistent positive specificity gain (Delta_spec = +1.18 pp) across all random seeds under the canonical protocol, and auxiliary crossing-number prediction exhibits robust performance across data regimes without the real-versus-random reversal observed for centroid alignment. Causal probes reveal that background changes alone flip 17-32% of predictions and phone-photo accuracy drops by 58-69 percentage points, underscoring that appearance bias remains the principal obstacle to deployment. These results collectively demonstrate that our diagnostic workflow provides a principled and practical tool for evaluating whether a hand-crafted structural prior delivers genuine task-relevant benefit beyond generic regularization.

绳结分类拓扑感知结构先验模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。