通过正则化微调提升3D多模态模型跨域适应能力
Domain Generalizable Adaptation of 3D Vision-Language Models via Regularized Fine-Tuning

- 选择性层微调结合多视角一致性与文本多样性正则
- 在多个基准上提升跨域泛化性能1.36%~3.11%
- 适合数据少、需鲁棒跨域迁移的3D视觉任务
领域自适应仍是3D视觉的核心挑战,尤其对于对齐3D点云与视觉、文本数据的多模态基础模型。尽管这些模型具备强大泛化能力,但在下游领域数据有限时微调常导致过拟合和灾难性遗忘。为此,我们提出ReFine3D,一种面向3D大多模态模型(LMMs)的正则化微调框架,实现领域泛化性优化。ReFine3D结合选择性层微调与两项针对性正则策略:通过对点云增强后多视角一致性约束,以及利用大语言模型生成同义词提示以增强文本多样性。此外,引入点渲染视觉监督和基于置信度聚合的测试时增强机制,进一步提升鲁棒性。在多个3D领域泛化基准上的实验表明,ReFine3D在基类到新类泛化上提升1.36%,跨数据集迁移提升2.43%,抗噪声破坏能力提升1.80%,少样本准确率最高提升3.11%,优于现有最优方法,且计算开销极低。
原文摘要 · Abstract (English)
Domain adaptation remains a central challenge in 3D vision, especially for multimodal foundation models that align 3D point clouds with visual and textual data. While these models demonstrate strong general capabilities, adapting them to downstream domains with limited data often leads to overfitting and catastrophic forgetting. To address this, we introduce ReFine3D, a regularized fine-tuning framework designed for domain-generalizable tuning of 3D large multimodal models (LMMs). ReFine3D combines selective layer tuning with two targeted regularization strategies: multi-view consistency across augmented point clouds and text diversity through synonym-based prompts generated by large language models. Additionally, we incorporate point-rendered vision supervision and a test-time augmentation mechanism with confidence-based aggregation to further enhance robustness. Extensive experiments across different 3D domain generalization benchmarks show that ReFine3D improves base-to-novel class generalization by 1.36%, cross-dataset transfer by 2.43%, robustness to corruption by 1.80%, and few-shot accuracy by up to 3.11%, outperforming prior state-of-the-art methods with minimal added computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。