arXiv:2509.16268cs.SEcs.AI2025-09

用因果分析揭示大模型函数调用如何提升指令遵循与安全检测能力

Digging Into the Internal: Causality-Based Analysis of LLM Function Calling

  • 通过层级与词元级因果干预,解析函数调用对模型内部逻辑的影响
  • 函数调用使恶意输入检测准确率平均提升135%,显著优于传统提示方法
  • 为提升大模型安全性与可控性提供可解释的机制支持,适合安全研究者参考

函数调用(FC)已成为大语言模型(LLMs)与外部系统交互、执行结构化任务的重要技术。然而其影响模型行为的内在机制仍不清晰。我们发现,除了常规用途,FC还能显著提升模型对用户指令的遵循程度。为此,我们引入因果分析方法,通过层级和词元级因果干预,剖析FC在响应用户查询时对模型内部计算逻辑的影响。分析结果证实了FC的显著作用,并揭示其深层机制。为验证结论,我们在两个基准数据集上对四种主流大模型进行了广泛实验,聚焦于提升模型安全鲁棒性这一关键应用场景。结果显示,相较于传统提示方法,基于FC的指令在检测恶意输入方面平均性能提升约135%,展现出增强大模型可靠性与实用能力的巨大潜力。

原文摘要 · Abstract (English)

Function calling (FC) has emerged as a powerful technique for facilitating large language models (LLMs) to interact with external systems and perform structured tasks. However, the mechanisms through which it influences model behavior remain largely under-explored. Besides, we discover that in addition to the regular usage of FC, this technique can substantially enhance the compliance of LLMs with user instructions. These observations motivate us to leverage causality, a canonical analysis method, to investigate how FC works within LLMs. In particular, we conduct layer-level and token-level causal interventions to dissect FC's impact on the model's internal computational logic when responding to user queries. Our analysis confirms the substantial influence of FC and reveals several in-depth insights into its mechanisms. To further validate our findings, we conduct extensive experiments comparing the effectiveness of FC-based instructions against conventional prompting methods. We focus on enhancing LLM safety robustness, a critical LLM application scenario, and evaluate four mainstream LLMs across two benchmark datasets. The results are striking: FC shows an average performance improvement of around 135% over conventional prompting methods in detecting malicious inputs, demonstrating its promising potential to enhance LLM reliability and capability in practical applications.

函数调用因果分析大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。