arXiv:2601.15549cs.CVcs.AI2026-01被引 2

用极少标注实现视频模型快速适应新场景,适合专家少的工业和医疗环境。

VIOLA: Towards Video In-Context Learning with Minimal Annotations

  • 通过密度与不确定性加权采样,精准挑选最有价值的少量样本。
  • 在仅需10%标注数据下,性能超越基线方法,达到稳定适配效果。
  • 专为标注稀缺场景设计,特别适合手术、工业等专业领域应用。

将多模态大语言模型(MLLMs)泛化到新视频领域对实际部署至关重要,但受限于标注数据稀缺。尽管上下文学习(ICL)提供免训练适配路径,传统方法依赖大量标注数据,而在工业或手术等专业场景中难以实现。为此,我们提出VIOLA(Video In-context Learning with Minimal Annotation),一种标签高效的框架,结合极少量专家标注与大量未标注数据。首先,提出密度-不确定性加权采样策略,避免视觉异常点,高效筛选兼具多样性、代表性与信息量的样本。其次,构建混合样本池,引入置信度感知检索与置信度感知提示机制,显式建模标签可靠性,基于相似性与置信度复合得分检索示范样本,并使MLLM能自适应区分真实标签与噪声伪标签。在九个不同基准上使用四个MLLM的实验表明,本框架在低资源设置下显著优于多种基线,在仅需少量标注时仍实现鲁棒适应。

原文摘要 · Abstract (English)

Generalizing Multimodal Large Language Models (MLLMs) to novel video domains is essential for real-world deployment but remains challenging due to the scarcity of labeled data. While In-Context Learning (ICL) offers a training-free adaptation path, standard methods rely on large annotated pools, which are often impractical in specialized environments like industrial or surgical settings since they require the experts' annotations. To bridge this gap, we introduce VIOLA (Video In-cOntext Learning with minimal Annotation), a label-efficient framework that synergizes minimal expert supervision with abundant unlabeled data. First, to maximize the efficiency of a strict annotation budget, we propose density-uncertainty-weighted sampling. Unlike standard diversity or uncertainty strategies that risk selecting visual outliers, our method leverages density estimation to identify samples that are simultaneously diverse, representative, and informative. Second, to utilize the remaining unlabeled data without noise propagation, we construct a hybrid pool and introduce confidence-aware retrieval and confidence-aware prompting. These mechanisms explicitly model label reliability, retrieving demonstrations based on a composite score of similarity and confidence while enabling the MLLM to adaptively distinguish between verified ground truths and noisy pseudo-labels. Extensive experiments across nine diverse benchmarks using four MLLMs demonstrate that our framework significantly outperforms various baselines in low-resource settings, achieving robust adaptation with minimal annotation costs.

视频理解上下文学习少样本学习多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。