用分阶段低秩适配融合多模态数据,提升医疗训练场景动作识别效果。
LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments
- 分步融合不同模态,先合相关性强的,再加异构模态。
- 在两个医疗训练数据集上表现优于单一模态模型。
- 参数高效,适合不同模态组合的数据集扩展。
本文提出一种基于低秩适配(LoRA)的级联多模态融合框架,用于医疗训练环境中的动作与活动识别。该架构结合参数高效的模态专属适配与顺序融合机制,使各模态可分阶段集成,且无需重新训练已有组件。不同于固定融合结构,本方法优先融合相关性较高的模态,再逐步引入异构模态,支持在具有不同模态组合的数据集间可扩展适应。我们在两个面向医疗训练环境的数据集——NurViD 和 Nurse Training dataset——上进行了评估。初步结果表明,所提出的级联融合策略优于单模态模型,并在性能上达到此前针对特定数据集报告基线的水平。整体结果表明,级联式LoRA融合是一种在医疗训练动作识别任务中整合异构模态的有前景的参数高效方法。
原文摘要 · Abstract (English)
This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific adaptation with sequential fusion, enabling modalities to be integrated in stages without retraining previously learned components. Rather than assuming a fixed fusion structure, the framework first integrates more closely related modalities and then incorporates additional heterogeneous modalities, supporting scalable adaptation across datasets with different modality sets.We evaluate the framework on two healthcare-oriented training environment datasets: NurViD and the Nurse Training dataset. Across these datasets, preliminary results suggest that the proposed cascaded fusion strategy improves over individual modality models and provides competitive performance relative to previously reported dataset-specific baselines. Overall, these findings indicate that cascaded LoRA-based fusion is a promising parameter-efficient approach for integrating heterogeneous modalities in medical training action and activity recognition tasks. github: https://github.com/anonymous0-ai/LoRA-Based-Cascaded-Multimodal-Fusion-.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。