自动构建企业级文本转Cypher查询的可演进基准测试
PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems

- 基于真实图谱和用户提问生成可执行的自然语言-查询对
- 产出3000个平衡且可验证的金融/社交网络测试用例
- 适合企业落地Text2Cypher系统的模型评估与优化
企业属性图在模式结构、术语体系、领域假设、治理约束和用户交互模式上差异显著。一个部署相关的文本转Cypher基准应反映用户和代理对该图实际提出的疑问。由于模式和数据唯一且随时间变化,创建此类基准极具挑战:每对自然语言查询与Cypher语句必须可执行、使用真实图实体、保持多样性,并在查询类型和难度上均衡。我们提出PIPE-Cypher,一种本地基准生成流水线,将实时属性图及可选的种子查询(来自客户问题、分析师日志或代理工具调用)转化为平衡的文本到Cypher基准。PIPE-Cypher结合模式分析、反向查询定位、受控生成、确定性Cypher治理、执行验证、脱敏处理、多样性控制及校准后的本地大模型判断器。采用本地Qwen3.5-9B进行生成与评判,输出3000个通过验证的FinBench/SNB示例,完成三次审计消融实验,以人工标签校准判断器行为,并评估11个本地下游模型。结果基准具有明确区分度:零样本迁移效果差,而少量样本提示显示,特定模式的示例库能有效提升兼容模型族性能。PIPE-Cypher使文本转Cypher基准测试成为可重复、随图谱、用户和工作负载演进的过程。
原文摘要 · Abstract (English)
Enterprise property graphs vary widely in schema structure, internal terminology, domain assumptions, governance constraints, and user interaction patterns. A deployment-relevant Text2Cypher benchmark therefore reflects the questions users and agents actually ask of that graph. Creating such a benchmark is difficult because schemas and values are unique, and graph structure changes over time. Each NL-query pair must also be executable, use real graph entities, preserve diversity, and remain balanced across query types and difficulty levels. We present PIPE-Cypher, a local benchmark-generation pipeline that turns a live property graph and optional seed queries from customer questions, analyst logs, or agent tool calls into balanced NL-to-Cypher benchmarks. PIPE-Cypher combines schema profiling, reverse-query grounding, constrained generation, deterministic Cypher governance, execution validation, redaction, diversity controls, and a calibrated local LLM judge. Using local Qwen3.5-9B generation and judging, PIPE-Cypher exports 3,000 accepted FinBench/SNB examples, completes three audited ablation suites, calibrates judge behavior with human labels, and evaluates 11 local downstream models. The resulting benchmark is deliberately discriminative: zero-shot transfer is weak, while a few-shot control shows that schema-specific example banks can help compatible model families. Together, PIPE-Cypher makes Text2Cypher benchmarking a repeatable process that evolves with the graph, its users, and its target workloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。