arXiv:2512.20501cs.CV2025-12

让机器理解空间语言、医学文本和知识图谱,提升多模态识别能力。

Bridging Modalities and Transferring Knowledge: Enhanced Multimodal Understanding and Recognition

  • 用空间推理BERT将文字描述转为2D图形布局。
  • 医学文本可精准定位解剖图谱中的3D位置,导航更准确。
  • 支持无模态计算的轻量级动作识别,适合实际部署。

本文探讨多模态对齐、翻译、融合与迁移,以增强机器对复杂输入的理解。第三章提出空间推理BERT,将文本中的空间关系转化为剪贴画的二维布局,实现人类空间认知一致的场景自动生成。第四章提出一种将医学文本映射到解剖图谱特定3D位置的方法,通过利用医学术语的空间共现设计损失函数,显著提升文本导航可解释性。第五章研究结构化文本向知识图谱中规范事实的转换,构建基准数据集以解决自然语言提取中的歧义问题,提供更清晰的可操作洞察。第六章提出融合视频帧与目标检测表示的多模态动作识别方法,提升识别鲁棒性与准确性。第七章探索用于第一人称动作识别的多模态知识迁移,证明多模态知识蒸馏可使仅含RGB的模型模仿多模态融合性能,降低计算开销同时保持效果。这些工作推进了空间语言理解、医学文本解析、知识图谱增强与动作识别的算法发展,提升系统处理多样化多模态输入的能力。

原文摘要 · Abstract (English)

This manuscript explores multimodal alignment, translation, fusion, and transference to enhance machine understanding of complex inputs. We organize the work into five chapters, each addressing unique challenges in multimodal machine learning. Chapter 3 introduces Spatial-Reasoning Bert for translating text-based spatial relations into 2D arrangements between clip-arts. This enables effective decoding of spatial language into visual representations, paving the way for automated scene generation aligned with human spatial understanding. Chapter 4 presents a method for translating medical texts into specific 3D locations within an anatomical atlas. We introduce a loss function leveraging spatial co-occurrences of medical terms to create interpretable mappings, significantly enhancing medical text navigability. Chapter 5 tackles translating structured text into canonical facts within knowledge graphs. We develop a benchmark for linking natural language to entities and predicates, addressing ambiguities in text extraction to provide clearer, actionable insights. Chapter 6 explores multimodal fusion methods for compositional action recognition. We propose a method fusing video frames and object detection representations, improving recognition robustness and accuracy. Chapter 7 investigates multimodal knowledge transference for egocentric action recognition. We demonstrate how multimodal knowledge distillation enables RGB-only models to mimic multimodal fusion-based capabilities, reducing computational requirements while maintaining performance. These contributions advance methodologies for spatial language understanding, medical text interpretation, knowledge graph enrichment, and action recognition, enhancing computational systems' ability to process complex, multimodal inputs across diverse applications.

多模态知识蒸馏空间推理医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。