为大模型辅助的振动能量采集器设计提供分阶段诊断基准
VEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester Design

- 构建763个基于文献的任务,用物理原理验证器评分
- 发现大模型在不同设计阶段表现差异显著,无单一模型全程领先
- 适合评估和优化工程场景下大模型的流程适配能力
无电池物联网需要在耦合物理约束下迭代设计振动能量采集器(VEH),而大模型正成为工程工作流的接口层。然而,现有工程基准主要评估最终成果有效性,难以揭示大模型在多阶段耦合设计中的行为特征。本文提出VEHBench,一个面向大模型辅助VEH设计的工程原生诊断基准,包含763个基于文献的任务,由分析型物理验证器评分。该基准评估四种设计角色:需求筛选、验证引导搜索、受损状态恢复、策略条件选择。实验表明,大模型能力显著依赖设计阶段:无单一模型在全流程中持续领先,响应控制特性揭示了各角色下的独特行为模式。因此,VEHBench为评估、选型、路由与改进验证器驱动的工程大模型提供了阶段感知基础。基准数据集已公开于https://huggingface.co/datasets/AnonymousVehbench/vehbench。
原文摘要 · Abstract (English)
Battery-free Internet of Things (IoT) requires iterative design of vibration energy harvesters (VEHs) under coupled physical constraints, while LLMs are emerging as interface layers for engineering workflows. However, existing engineering benchmarks primarily assess final artifact validity, offering limited insights into how LLMs behave across different stages of coupled physical design. We introduce VEHBench, an engineering-native diagnostic benchmark for LLM-assisted VEH design, featuring 763 literature-grounded tasks scored by an analytical physical oracle. VEHBench evaluates four design roles: specification triage, verifier-guided search, corrupted-state recovery, and policy-conditioned selection. Experimental results reveal that LLM capability is strongly stage-dependent: no single model consistently dominates the entire workflow, and response-control profiles expose distinct behavioral patterns across design roles. VEHBench thus provides a stage-aware foundation for evaluating, selecting, routing, and improving verifier-grounded engineering LLMs. The benchmark artifact is available at https://huggingface.co/datasets/AnonymousVehbench/vehbench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。