ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
A new benchmark called ESQ‑Bench has been released to evaluate natural‑language‑to‑SQL (NL2SQL) systems on enterprise‑grade Oracle databases, addressing the gap between academic testbeds and real‑world complexity. The benchmark comprises six fully populated schemas—totaling 465 tables and 164,682 rows—with identical seed data replicated on Oracle, PostgreSQL, MySQL, and SQL Server. It offers 550 gold‑validated natural‑language questions paired with correct SQL queries, divided across three tiers of schema complexity (Tier‑1: 95 pairs, Tier‑2: 228 pairs, Tier‑3: 227 pairs). Evaluation uses a four‑metric harness measuring exact match (EM), execution accuracy (EX), semantic relevance (SR), and silent divergence (SD), allowing researchers to assess both surface‑level correctness and deeper semantic fidelity.
Testing with leading closed‑API models reveals a sharp decline in performance as schema complexity rises. When prompts are linked to the schema, GPT‑4o’s execution accuracy drops from 79.8 % on Tier‑1 to 60.3 % on Tier‑2 and 57.2 % on Tier‑3, while its exact‑match rate remains below 7 % across all tiers. Silent‑divergence—cases where a query executes without error but returns incorrect results—reaches between 73 % and 99 % among the queries that do execute, indicating that most failures are semantic rather than syntactic. Claude Sonnet 4.6 outperforms GPT‑4o on every tier with execution accuracies of 87.4 %, 74.9 %, and 68.7 % respectively, whereas zero‑shot prompting with GPT‑4o reverses the tier trend, achieving higher execution rates on the harder tiers due to survivor bias. An open‑weight baseline, Llama 3.2, manages only 13.3 % execution accuracy across the full benchmark, underscoring the difficulty of enterprise‑level NL2SQL tasks for locally run models.
The findings highlight a substantial gap between current state‑of‑the‑art NL2SQL systems and the demands of production database environments, where silent semantic errors can be especially costly. The high silent‑divergence rates suggest that models may appear successful under traditional metrics while delivering incorrect query results, a risk that escalates with schema intricacy. By providing a multi‑tier, Oracle‑first benchmark and a comprehensive evaluation suite, ESQ‑Bench offers the research community a realistic testing ground to drive improvements in model robustness, prompting strategies, and cross‑dialect compatibility, ultimately aiming to bridge the divide between academic progress and enterprise applicability.
⚡ Effects Interpreter
🌍World Economy
- ▶Currencies can shift on this kind of news, changing the cost of holidays and imports.
- ▶Markets typically move first and ask questions later after headlines like these.
🏙️Local Economy
- ▶Community shops may gradually reprice stock as costs shift upstream.
- ▶Everyday costs in your town may drift as the wider economy reacts.
🏦Rates & Banks
- ▶Banks generally prefer a wait-and-see approach before touching their rates.
- ▶Lenders often hold their nerve until a clearer trend appears.
❤️Health
- ▶Being kind to yourself matters just as much as staying informed.
- ▶Community wellbeing might dip a touch while people wait for clarity.
💷Wealth
- ▶Short-term swings are normal — long-term savers usually stay the course.
- ▶Day-to-day swings rarely matter much if your money is invested for years.
🏠Housing
- ▶Rental yields in the area may shift only slightly, if at all, from this.
- ▶Renters may see costs drift a touch as landlords weigh their own bills.