arXiv:2512.08079cs.IR2025-12

用AI自动分组并描述法律取证中的海量图像,省时省力。

Leveraging Machine Learning and Large Language Models for Automated Image Clustering and Description in Legal Discovery

  • 用K-means将图像分为20个视觉相似组,选20张代表图生成基础描述。
  • 分组描述用LLM效果最好,比传统方法更准确且覆盖更全。
  • 按分布采样20张图就足够,比全量处理快得多,适合大规模应用。

数字图像的快速增长给法律取证、数字存档和内容管理带来巨大挑战。企业与法律团队需在紧迫时间内对大规模图像集进行组织、分析并提取有效信息,手动审查已不现实且成本高昂。为此,本文系统研究了融合图像聚类、图像描述生成与大语言模型(LLM)的自动化聚类描述生成方法。我们采用K-means将图像划分为20个视觉一致的簇,并利用Azure AI Vision API生成基础描述。评估三个关键维度:(1)采样策略,对比随机、质心、分层、混合与密度采样与全量使用簇内图像的效果;(2)提示技术,比较标准提示与思维链提示;(3)描述生成方法,对比基于LLM的方法与传统TF-IDF及模板法。通过语义相似性与覆盖率评估质量。结果表明,每簇选取20张图像的策略在性能上接近全量使用,显著降低计算成本,仅分层采样略有下降。基于LLM的方法始终优于TF-IDF基线,标准提示优于思维链提示。研究为法律取证等高吞吐场景提供了可扩展、高精度的自动化图像组织方案。

原文摘要 · Abstract (English)

The rapid increase in digital image creation and retention presents substantial challenges during legal discovery, digital archive, and content management. Corporations and legal teams must organize, analyze, and extract meaningful insights from large image collections under strict time pressures, making manual review impractical and costly. These demands have intensified interest in automated methods that can efficiently organize and describe large-scale image datasets. This paper presents a systematic investigation of automated cluster description generation through the integration of image clustering, image captioning, and large language models (LLMs). We apply K-means clustering to group images into 20 visually coherent clusters and generate base captions using the Azure AI Vision API. We then evaluate three critical dimensions of the cluster description process: (1) image sampling strategies, comparing random, centroid-based, stratified, hybrid, and density-based sampling against using all cluster images; (2) prompting techniques, contrasting standard prompting with chain-of-thought prompting; and (3) description generation methods, comparing LLM-based generation with traditional TF-IDF and template-based approaches. We assess description quality using semantic similarity and coverage metrics. Results show that strategic sampling with 20 images per cluster performs comparably to exhaustive inclusion while significantly reducing computational cost, with only stratified sampling showing modest degradation. LLM-based methods consistently outperform TF-IDF baselines, and standard prompts outperform chain-of-thought prompts for this task. These findings provide practical guidance for deploying scalable, accurate cluster description systems that support high-volume workflows in legal discovery and other domains requiring automated organization of large image collections.

图像聚类法律AI大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。