arXiv:2506.12116cs.CLcs.AI2025-06被引 4

用多模态嵌入无监督聚类文档模板与类别,提升识别鲁棒性。

Unsupervised Document and Template Clustering using Multimodal Embeddings

  • 融合文本、版式、视觉编码器输出,生成分类型文档向量
  • 视觉特征在清晰文档上几乎解决模板发现,文本主导噪声场景
  • 提供可复现的调参方案,适合做无监督文档组织的研究者

我们研究使用冻结的多模态编码器与经典聚类算法,在无监督条件下对文档进行类别和模板层面的聚类。提出一种模型无关的流程:(i) 将文本-版式-视觉编码器的最后层状态投影为带标记类型的文档向量;(ii) 使用质心或密度基方法(包括 HDBSCAN + k-NN 去除未分类点)进行聚类。在五个数据集上评估八种编码器(纯文本、版式感知、纯视觉、视觉语言)与四种聚类方法(k-Means、DBSCAN、HDBSCAN + k-NN、BIRCH),涵盖干净合成发票、严重降质的打印扫描件、扫描收据及真实身份与证书文档。结果揭示了不同模态的失效模式与准确率-鲁棒性权衡:视觉特征在清晰页面上近乎完全解决模板发现,而文本在协变量偏移下占优;融合编码器表现最佳。本文还提供了可复现的无标签调参协议与标准化评估设置,以指导未来无监督文档组织研究。

原文摘要 · Abstract (English)

We study unsupervised clustering of documents at both the category and template levels using frozen multimodal encoders and classical clustering algorithms. We systematize a model-agnostic pipeline that (i) projects heterogeneous last-layer states from text-layout-vision encoders into token-type-aware document vectors and (ii) performs clustering with centroid- or density-based methods, including an HDBSCAN + $k$-NN assignment to eliminate unlabeled points. We evaluate eight encoders (text-only, layout-aware, vision-only, and vision-language) with $k$-Means, DBSCAN, HDBSCAN + $k$-NN, and BIRCH on five corpora spanning clean synthetic invoices, their heavily degraded print-and-scan counterparts, scanned receipts, and real identity and certificate documents. The study reveals modality-specific failure modes and a robustness-accuracy trade-off, with vision features nearly solving template discovery on clean pages while text dominates under covariate shift, and fused encoders offering the best balance. We detail a reproducible, oracle-free tuning protocol and the curated evaluation settings to guide future work on unsupervised document organization.

无监督学习文档分析多模态嵌入聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。