arXiv:2507.16844cs.LG2025-07

用视觉语言模型读懂第三方时序图,提升设计验证效率

TD-Interpreter: Enhancing the Understanding of Timing Diagrams with Visual-Language Learning

  • 基于微调的70亿参数多模态模型,理解时序图与文本问题
  • 合成数据训练使模型在基准测试中显著超越未调优的GPT-4o
  • 适合硬件设计与验证工程师快速解析复杂时序图

我们提出TD-Interpreter,一种专为工程师设计的ML工具,用于理解第三方提供的复杂时序图(TDs),辅助其设计与验证流程。该工具构建于多模态问答环境,允许用户输入一组时序图并提出相关设计与验证问题。我们采用微调后的轻量级7B多模态大语言模型(LLaVA)实现该系统,并针对训练数据稀缺问题,开发了将视觉信息与文本解释对齐的合成数据生成流程。实验评估表明,TD-Interpreter在多个基准测试中表现优异,显著优于未调优的GPT-4o。

原文摘要 · Abstract (English)

We introduce TD-Interpreter, a specialized ML tool that assists engineers in understanding complex timing diagrams (TDs), originating from a third party, during their design and verification process. TD-Interpreter is a visual question-answer environment which allows engineers to input a set of TDs and ask design and verification queries regarding these TDs. We implemented TD-Interpreter with multimodal learning by fine-tuning LLaVA, a lightweight 7B Multimodal Large Language Model (MLLM). To address limited training data availability, we developed a synthetic data generation workflow that aligns visual information with its textual interpretation. Our experimental evaluation demonstrates the usefulness of TD-Interpreter which outperformed untuned GPT-4o by a large margin on the evaluated benchmarks.

时序图多模态LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。