用图文预测真菌生长阶段和时间,让模型学会看变化
CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text
- 基于CLIP架构,联合学习图像文本嵌入并预测时间
- 在合成数据集上同时准确识别生长阶段与连续时间戳
- 适合生物监测、动态过程分析等需要时间感知的应用
理解生物生长的时序动态在微生物学、农业和生物降解研究中至关重要。尽管对比语言图像预训练(CLIP)等视觉-语言模型在跨模态推理方面表现优异,但其对时序进展的捕捉能力有限。为此,我们提出CLIPTime,一种多模态、多任务框架,可从图像和文本输入中预测真菌生长的发育阶段及对应时间戳。该模型基于CLIP架构,学习联合视觉-文本嵌入,实现无需显式时间输入的时序感知推理。为支持训练与评估,我们构建了一个带有对齐时间戳和分类阶段标签的合成真菌生长数据集。CLIPTime联合执行分类与回归任务,同时预测离散生长阶段和连续时间。我们还设计了定制化评估指标,包括时序准确率与回归误差,以衡量时间感知预测的精度。实验表明,CLIPTime能有效建模生物演化过程,并生成可解释、时序锚定的输出,凸显视觉-语言模型在真实世界生物监测中的潜力。
原文摘要 · Abstract (English)
Understanding the temporal dynamics of biological growth is critical across diverse fields such as microbiology, agriculture, and biodegradation research. Although vision-language models like Contrastive Language Image Pretraining (CLIP) have shown strong capabilities in joint visual-textual reasoning, their effectiveness in capturing temporal progression remains limited. To address this, we propose CLIPTime, a multimodal, multitask framework designed to predict both the developmental stage and the corresponding timestamp of fungal growth from image and text inputs. Built upon the CLIP architecture, our model learns joint visual-textual embeddings and enables time-aware inference without requiring explicit temporal input during testing. To facilitate training and evaluation, we introduce a synthetic fungal growth dataset annotated with aligned timestamps and categorical stage labels. CLIPTime jointly performs classification and regression, predicting discrete growth stages alongside continuous timestamps. We also propose custom evaluation metrics, including temporal accuracy and regression error, to assess the precision of time-aware predictions. Experimental results demonstrate that CLIPTime effectively models biological progression and produces interpretable, temporally grounded outputs, highlighting the potential of vision-language models in real-world biological monitoring applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。