ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence

A new benchmark called ESQ‑Bench has been released to evaluate natural‑language‑to‑SQL (NL2SQL) systems on enterprise‑grade Oracle databases, addressing the gap between academic testbeds and real‑world complexity. The benchmark comprises six fully populated schemas—totaling 465 tables and 164,682 rows—with identical seed data replicated on Oracle, PostgreSQL, MySQL, and SQL Server. It offers 550 gold‑validated natural‑language questions paired with correct SQL queries, divided across three tiers of schema complexity (Tier‑1: 95 pairs, Tier‑2: 228 pairs, Tier‑3: 227 pairs). Evaluation uses a four‑metric harness measuring exact match (EM), execution accuracy (EX), semantic relevance (SR), and silent divergence (SD), allowing researchers to assess both surface‑level correctness and deeper semantic fidelity.

Testing with leading closed‑API models reveals a sharp decline in performance as schema complexity rises. When prompts are linked to the schema, GPT‑4o’s execution accuracy drops from 79.8 % on Tier‑1 to 60.3 % on Tier‑2 and 57.2 % on Tier‑3, while its exact‑match rate remains below 7 % across all tiers. Silent‑divergence—cases where a query executes without error but returns incorrect results—reaches between 73 % and 99 % among the queries that do execute, indicating that most failures are semantic rather than syntactic. Claude Sonnet 4.6 outperforms GPT‑4o on every tier with execution accuracies of 87.4 %, 74.9 %, and 68.7 % respectively, whereas zero‑shot prompting with GPT‑4o reverses the tier trend, achieving higher execution rates on the harder tiers due to survivor bias. An open‑weight baseline, Llama 3.2, manages only 13.3 % execution accuracy across the full benchmark, underscoring the difficulty of enterprise‑level NL2SQL tasks for locally run models.

The findings highlight a substantial gap between current state‑of‑the‑art NL2SQL systems and the demands of production database environments, where silent semantic errors can be especially costly. The high silent‑divergence rates suggest that models may appear successful under traditional metrics while delivering incorrect query results, a risk that escalates with schema intricacy. By providing a multi‑tier, Oracle‑first benchmark and a comprehensive evaluation suite, ESQ‑Bench offers the research community a realistic testing ground to drive improvements in model robustness, prompting strategies, and cross‑dialect compatibility, ultimately aiming to bridge the divide between academic progress and enterprise applicability.

Sources cited: 📰 ArXiv AI ↗

⚡ Effects Interpreter

🌍World Economy

  • Confidence among international firms may wobble until the picture clears.
  • Cross-border money flows can gradually change direction after events like this.

🏙️Local Economy

  • The high street usually mirrors big-picture shifts, just a little later.
  • Household budgets may notice a small ripple in due course.

🏦Rates & Banks

  • Borrowing costs might hold steady for now, but they can turn on fresh news.
  • Your loan or mortgage rate is more likely to drift than to lurch here.

❤️Health

  • Local health services could get busier depending on how things develop.
  • Unsettling news can weigh on sleep and mood, so peace of mind matters.

💷Wealth

  • Savings and portfolios can see short-lived ups and downs after news like this.
  • It may be worth a quick look at your ISA or pension in the coming days.

🏠Housing

  • Mortgage deals could edge around if lenders read the wider mood.
  • Buyers and renters might notice only a gentle drift, if anything at all.
Share: 𝕏 Twitter Facebook LinkedIn WhatsApp

Editorial note: This analysis was produced by the News Effects Interpreter, an AI editorial tool that cross-references 1 independent news sources and contextualises events in terms of their real-world impact on ordinary people. Original reporting is linked above. News Effects does not alter the facts of source reports.