arXiv:2512.08337cs.CV2025-12

用T1图像生成功能脑图,提升脑成像修复效率

DINO-BOLDNet: A DINOv3-Guided Multi-Slice Attention Network for T1-to-BOLD Generation

  • 结合DINOv3结构特征与多切片注意力融合上下文
  • 在248人临床数据上PSNR和MS-SSIM均优于基线
  • 首个直接从T1生成平均BOLD图的框架,适合脑影像重建

从T1加权图像生成BOLD图像为恢复缺失或受损的BOLD信息提供了有效途径,支持下游分析任务。为此,我们提出DINO-BOLDNet,一种基于DINOv3引导的多切片注意力网络,将冻结的自监督DINOv3编码器与轻量级可训练解码器结合。模型利用DINOv3提取单切片结构特征,并通过独立的切片注意力模块融合相邻切片的上下文信息。多尺度生成解码器恢复精细的功能对比度,同时基于DINO的感知损失在Transformer特征空间中保持预测与真实BOLD在结构和纹理上的一致性。在包含248名受试者的临床数据集上实验表明,DINO-BOLDNet在PSNR和MS-SSIM指标上均优于条件GAN基线。据我们所知,这是首个能够直接从T1w图像生成均值BOLD图像的框架,凸显了自监督Transformer引导在结构到功能映射中的潜力。

原文摘要 · Abstract (English)

Generating BOLD images from T1w images offers a promising solution for recovering missing BOLD information and enabling downstream tasks when BOLD images are corrupted or unavailable. Motivated by this, we propose DINO-BOLDNet, a DINOv3-guided multi-slice attention framework that integrates a frozen self-supervised DINOv3 encoder with a lightweight trainable decoder. The model uses DINOv3 to extract within-slice structural representations, and a separate slice-attention module to fuse contextual information across neighboring slices. A multi-scale generation decoder then restores fine-grained functional contrast, while a DINO-based perceptual loss encourages structural and textural consistency between predictions and ground-truth BOLD in the transformer feature space. Experiments on a clinical dataset of 248 subjects show that DINO-BOLDNet surpasses a conditional GAN baseline in both PSNR and MS-SSIM. To our knowledge, this is the first framework capable of generating mean BOLD images directly from T1w images, highlighting the potential of self-supervised transformer guidance for structural-to-functional mapping.

脑影像生成自监督学习多切片建模DINOv3

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。