arXiv:2410.10741cs.AIcs.LG2024-10被引 8

用真实传感器数据测试大模型编程处理能力,发现其在复杂任务上仍逊于专家。

SensorBench: Benchmarking LLMs in Coding-Based Sensor Processing

  • 构建包含多种真实传感器数据的评测基准SensorBench
  • 大模型在简单任务中表现良好,但复杂参数任务仍不及工程专家
  • 自验证提示策略在48%任务中优于其他方法,适合开发传感器系统

传感器数据的有效处理、解读与管理已成为网络物理系统的关键环节。传统上,该过程需要深厚的理论知识和信号处理工具熟练度。近期研究表明,大语言模型(LLMs)在处理传感数据方面展现出潜力,或可作为传感系统开发的智能协作者。为探索这一潜力,我们构建了综合性基准SensorBench,以建立可量化的评估目标。该基准涵盖多种真实世界传感器数据集,用于不同任务。结果表明,尽管大模型在简单任务中表现出色,但在涉及参数选择的组合式任务中,仍面临固有挑战,远不及工程专家。此外,我们研究了四种提示策略,发现自验证策略在48%的任务中优于所有基线。本研究为未来基于大模型的传感器处理提供了全面基准与提示分析,推动构建智能化传感系统协作者。

原文摘要 · Abstract (English)

Effective processing, interpretation, and management of sensor data have emerged as a critical component of cyber-physical systems. Traditionally, processing sensor data requires profound theoretical knowledge and proficiency in signal-processing tools. However, recent works show that Large Language Models (LLMs) have promising capabilities in processing sensory data, suggesting their potential as copilots for developing sensing systems. To explore this potential, we construct a comprehensive benchmark, SensorBench, to establish a quantifiable objective. The benchmark incorporates diverse real-world sensor datasets for various tasks. The results show that while LLMs exhibit considerable proficiency in simpler tasks, they face inherent challenges in processing compositional tasks with parameter selections compared to engineering experts. Additionally, we investigate four prompting strategies for sensor processing and show that self-verification can outperform all other baselines in 48% of tasks. Our study provides a comprehensive benchmark and prompting analysis for future developments, paving the way toward an LLM-based sensor processing copilot.

大模型传感器编程评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。