arXiv:2607.08798cs.GRcs.CV2026-07

用现成大模型零样本识别应力关键区域,提升3D网格细化精度。

GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement

论文配图:GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement
图 1 · 摘自论文原文
  • 通过物理提示让多模态大模型定位应力敏感区
  • 在匹配预算下精度优于传统几何启发式方法
  • 适合需要快速自动化仿真的工程师和研究者

自适应体积有限元网格划分是计算机辅助工程与分析中的关键步骤,直接影响计算成本。传统方法依赖迭代偏微分方程求解器或需大量仿真数据训练的监督型代理模型。尽管多模态大语言模型(MLLMs)在2D视觉任务中表现优异,其零样本下基于几何理解与物理知识进行语义定位的能力仍未知。本文探索核心问题:预训练的MLLMs能否作为可行的零样本几何代理,用于有限元网格细化?为此,提出GReFEM框架,利用MLLMs根据物理引导文本提示,可视化定位应力关键区域。为弥合2D MLLM预训练与3D几何间的差距,引入orthoViews视图选择模块,最大化关键几何特征的可观测性。在多种CAD几何、载荷工况及主流MLLM上进行深入实证评估,对比经调优的几何启发式方法,在严格匹配的细化预算下表现。结果表明,MLLMs具备强零样本能力,能准确响应复杂时空物理指令,对应力相关特征的定位精度高于盲目启发式方法。本研究揭示了当前MLLMs在物理语义对齐上的成功与局限,定义了基础模型作为自动化仿真流程语义助手的新前沿。

原文摘要 · Abstract (English)

Adaptive volumetric finite element meshing is a critical step in computer-aided engineering and analysis that dictates the computational budget of a given problem. It traditionally requires iterative PDE solvers or heavily supervised, data-driven surrogates trained on large-scale simulation data. While Multimodal Large Language Models (MLLMs) excel in 2D visual tasks, their zero-shot capability to semantically ground regions based on geometric understanding and physics remains an open question. Overall, this study explores a significant question: can the high-level semantic understanding of off-the-shelf MLLMs serve as a viable, zero-shot geometric proxy for finite element mesh refinement? To investigate this, we introduce GReFEM (Geometric Reasoning Enhanced Multimodal LLMs for Finite Element Meshing), a framework that utilizes MLLMs to visually localize stress-critical regions based on physics-guided textual prompts. To bridge the gap between 2D MLLM pre-training and 3D geometries, we introduce orthoViews, a view-selection module that maximizes the observability of key geometric features. We conduct an in-depth empirical evaluation across diverse CAD geometries, loading cases, and SOTA MLLMs, comparing them against a tuned geometric heuristic under a strict, matched refinement budget. Our findings reveal that MLLMs demonstrate robust zero-shot capacity to accurately follow complex spatial-physical instructions, isolating stress-relevant features with higher precision than blind heuristics. By mapping both the successes and current limitations of MLLMs in physical grounding, this study defines the frontier of foundation models as semantic assistants in automated simulation workflows.

3D网格细化多模态大模型物理引导零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。