arXiv:2604.01554cs.CRcs.LG2026-04

构建真实多样的二进制函数相似性评测基准,揭示现有模型泛化能力短板。

EXHIB: A Benchmark for Realistic and Diverse Evaluation of Function Similarity in the Wild

  • 构建五个来自真实场景的评测数据集,覆盖不同类型的二进制差异。
  • 9个模型在固件和语义数据集上性能下降最高达30%,暴露泛化缺陷。
  • 适合软件安全、漏洞分析与恶意代码研究者参考评估模型真实表现。

二进制函数相似性检测(BFSD)是软件安全的核心问题,支撑漏洞分析、恶意软件分类和补丁溯源等任务。过去几十年虽有众多模型与工具涌现,但因缺乏全面通用的基准,研究者难以有效比较。现有数据集范围有限,多聚焦于特定变换或二进制类型,无法反映真实应用的多样性。本文提出EXHIB,一个包含五个从真实世界收集的数据集的基准,分别突出BFSD问题空间的不同方面。我们在EXHIB上评估了9个代表性的模型,覆盖多种BFSD范式,发现其在固件和语义数据集上的性能下降最高达30%,揭示显著的泛化差距。结果表明,对低中层二进制变化的鲁棒性无法推广至高层语义差异,暴露出当前BFSD评估实践的关键盲区。

原文摘要 · Abstract (English)

Binary Function Similarity Detection (BFSD) is a core problem in software security, supporting tasks such as vulnerability analysis, malware classification, and patch provenance. In the past few decades, numerous models and tools have been developed for this application; however, due to the lack of a comprehensive universal benchmark in this field, researchers have struggled to compare different models effectively. Existing datasets are limited in scope, often focusing on a narrow set of transformations or types of binaries, and fail to reflect the full diversity of real-world applications. We introduce EXHIB, a benchmark comprising five realistic datasets collected from the wild, each highlighting a distinct aspect of the BFSD problem space. We evaluate 9 representative models spanning multiple BFSD paradigms on EXHIB and observe performance degradations of up to 30% on firmware and semantic datasets compared to standard settings, revealing substantial generalization gaps. Our results show that robustness to low- and mid-level binary variations does not generalize to high-level semantic differences, underscoring a critical blind spot in current BFSD evaluation practices.

函数相似性软件安全二进制分析评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。