首个电池数字护照合规性数据集,助力欧盟新规落地
BatteryPass-12K: The First Dataset for the Novel Digital Battery Passport Conformance Task

- 基于真实样本合成数据集,构建电池护照合规分类新任务
- 思维型模型表现最佳,小模型反而超越部分大模型
- 适合关注电池合规、AI安全与政策技术落地的研究者
我们提出了一项全新的数字电池护照(DBP)合规性分类任务,并发布了首个公开基准数据集BatteryPass-12K,该数据集由真实试点样本合成。随着欧盟电池法规即将生效,当前尚无公开数据集。我们评估了22个语言模型(包括小模型、专家混合模型和密集大模型)在零样本推理下的表现,还进行了少样本推理与提示注入攻击的额外分析。结果表明:(1) 思维型模型表现最优,GPT-5.4在验证集和测试集上的平均F1分别为0.98(0.03)和0.71(0.22);(2) 少样本示例显著提升性能;(3) 当前前沿模型仍面临挑战;(4) 模型参数规模扩大并不必然带来性能提升,小模型表现优于部分大模型;(5) 提示注入攻击会降低模型性能。尽管BatteryPass-12K仅基于真实试点样本,但其可拓展至电池生命周期推理等新兴任务。数据集已以宽松许可协议(CC-BY-4.0)公开。
原文摘要 · Abstract (English)
We introduce a novel task of digital battery passport (DBP) conformance classification and introduce the first public benchmark for the task: BatteryPass-12K, created synthetically from real pilot samples. This is as the EU's battery regulation on DBPs comes into effect soon and there exists no public dataset. We evaluated 22 language models (LMs) in zero-shot inference, spanning small LMs (SLMs), mixture of experts (MoEs), and dense LLMs. We also conducted analysis, additional evaluations of few-shot inference and prompt-injection attacks to find that (1) Thinking models have the best performance (with GPT-5.4 scoring 0.98 (0.03) and 0.71 (0.22) on average as F1 (and confidence interval at 95%) on the validation and test sets, respectively), (2) few-shot examples improve performance significantly, (3) generally capable frontier models find the task challenging, (4) merely scaling model parameters does not necessarily lead to improved performance, as SLMs outperformed some LLMs, and (5) prompt-injection attacks degrade performance. We note that BatteryPass-12K, though limited to real pilot samples, may be useful for other known or emerging tasks in the battery domain, e.g. lifecycle reasoning. We publicly release the dataset under a permissive licence (CC-BY-4.0).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。