用新基准测试发现:科学任务中微调和上下文学习各有优劣。
Adapting General-Purpose Foundation Models for X-ray Ptychography in Low-Data Regimes
- 对比微调与上下文学习,按任务类型选择最佳适配策略。
- 视觉任务上两者互补,文本任务上大模型上下文学习更优。
- 为科学领域AI系统设计提供可复用的适配框架。
先进显微技术的工作流自动化是重要目标,通用基础模型如语言模型(LLMs)和视觉语言模型(VLMs)展现出巨大潜力。然而,将其适配到特定科学任务仍具挑战,最优领域适配策略尚不明确。为此,我们提出PtychoBench——一个用于全息分析的多模态、多任务基准。通过该基准,我们系统比较了两种专业化策略:监督微调(SFT)与上下文学习(ICL)。在数据稀缺条件下,评估了基于VLM的视觉伪影检测任务和基于LLM的文本参数推荐任务。结果表明,最优路径取决于任务类型:视觉任务中,微调模型结合上下文感知示例表现最佳(均值F1为0.728);文本任务中,大模型的上下文学习优于微调模型,峰值F1达0.847,超过强基线“超专家”微调模型(零样本F1为0.839)。我们还验证了上下文感知提示的有效性,并发现微调模型存在持续的上下文干扰现象。结果在强基线(包括GPT-4o和基于DINOv3的分类器)上得到验证,为科学领域人工智能提供了关键洞见:适配路径依赖任务模态,为构建更高效的科学代理系统提供清晰框架。
原文摘要 · Abstract (English)
The automation of workflows in advanced microscopy is a key goal where foundation models like Language Models (LLMs) and Vision-Language Models (VLMs) show great potential. However, adapting these general-purpose models for specialized scientific tasks is critical, and the optimal domain adaptation strategy is often unclear. To address this, we introduce PtychoBench, a new multi-modal, multi-task benchmark for ptychographic analysis. Using this benchmark, we systematically compare two specialization strategies: Supervised Fine-Tuning (SFT) and In-Context Learning (ICL). We evaluate these strategies on a visual artifact detection task with VLMs and a textual parameter recommendation task with LLMs in a data-scarce regime. Our findings reveal that the optimal specialization pathway is task-dependent. For the visual task, SFT and ICL are highly complementary, with a fine-tuned model guided by context-aware examples achieving the highest mean performance (Micro-F1 of 0.728). Conversely, for the textual task, ICL on a large base model is the superior strategy, reaching a peak Micro-F1 of 0.847 and outperforming a powerful "super-expert" SFT model (0-shot Micro-F1 of 0.839). We also confirm the superiority of context-aware prompting and identify a consistent contextual interference phenomenon in fine-tuned models. These results, benchmarked against strong baselines including GPT-4o and a DINOv3-based classifier, offer key observations for AI in science: the optimal specialization path in our benchmark is dependent on the task modality, offering a clear framework for developing more effective science-based agentic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。