Shipping a Text-to-SQL agent that doesn't hallucinate column names
The dangerous thing about a Text-to-SQL demo is how quickly it can look finished. You type a question in plain language. The model replies with a neat query. The result table appears. People around the screen nod because the shape of the interaction feels right, and shape is often enough to fool us for a while. Then you move from the demo database to the real one. That is when the hallucinations begin to feel personal. A column name that almost exists. A join that works syntactically but quietly changes the meaning of the answer. A filter slipped onto the wrong timestamp field. A confident paragraph explaining a result that should never have been trusted in the first place. None of these failures are loud. That is what makes them expensive.
When I was building a Text-to-SQL workflow for stock-analysis use cases, I realised very quickly that the problem was not getting the model to generate SQL. Models are already good enough at producing something that resembles SQL. The real problem was teaching the system to respect the database as a physical place with rules, scars, and history. So we stopped treating the agent like a solitary genius and started treating it like a careful analyst who had to earn the right to speak. We added schema retrieval before generation, not after. We pushed richer metadata into the prompt, especially the ugly business meanings that column names never tell you on their own. We used semantic search to narrow which tables the model was even allowed to think about. We kept a library of expert-crafted examples because style matters more when ambiguity is expensive. And when the model still drifted, we added a second pass whose job was not to sound smart but to be suspicious.
What schema linking actually does
For a long time I treated schema linking as a preprocessing chore — a thing you bolt on so the prompt is shorter. After a few months in production, I think of it as the chassis of the whole system. Spider (Yu et al. 2018) and BIRD (Li et al. 2023) both make this point quietly: the hardest examples are not the ones with elaborate SQL, they are the ones where the question lives at the boundary between two columns that look almost the same. "Closing price" and "last price" sound interchangeable until you remember that one is the official end-of-session number and the other is whatever the last tick happened to be — sometimes mid-session, sometimes stale.
In our setup the linker does three things in sequence. It embeds the question and ranks tables by similarity, then walks the foreign-key graph to pick neighbors that frequently get joined together, then re-scores candidate columns using short business-meaning notes that we wrote by hand for the ugly ones. The notes are the unglamorous part. They look like "close_price = official end-of-day, only valid after 14:45 ICT" and they cost a senior analyst a Friday afternoon to produce. They are also the single highest-leverage thing we ever shipped.
Why we made the agent willing to abstain
The first version of the agent always answered. It was friendly, it was fast, and it was occasionally confidently wrong in ways that were almost impossible to catch on review. We borrowed an idea that has become a quiet consensus in the agent literature — Anthropic's Building Effective Agents (2024) phrases it as keeping agents simple and giving them clean exits — and we added an explicit abstain branch. If the linker's top-1 score was below a threshold, or if two candidate columns were too close to call, the agent stopped trying. It said something close to "I am not confident which price you mean — did you want the official close, or the latest tick?"
That one change shifted how analysts felt about the system more than any model upgrade we ever made. People stopped describing it as "the AI thing" and started describing it as "the assistant." In our internal eval — roughly two hundred analyst questions held out from training — adding an abstain path cut wrong-column rates from low double-digits down to a handful of percent, illustrative numbers we plot below. Every percentage point we gave up in coverage we got back twice in trust.
Self-correction does its share of the work. We followed the spirit of DIN-SQL (Pourreza & Rafiei 2023) and DAIL-SQL (Gao et al. 2023) — decompose the problem, generate a candidate, then run a critique pass that has the schema in front of it. Reflexion (Shinn et al. 2024) is a useful mental model for that critique: do not just regenerate, write down why the previous attempt was wrong and let the next attempt see that note. Self-Consistency (Wang et al. 2022) is the cheap upgrade — sample a few queries, keep the one that survives a vote, and accept that diversity catches more silent errors than any single clever decode. None of these are silver bullets. Stacked together, they make the agent measurably less embarrassing.
A note on benchmarks vs production analysts
Spider and BIRD are honest benchmarks, and I owe them a lot. But a benchmark question is a polite question. It comes with a clean schema, a single intended answer, and no political weight. A real analyst question often arrives as half a sentence in Slack, references a dashboard from last quarter, and assumes you already know which subsidiary is being reported on. The hardest part of the job is not the SQL — it is the silent context that the question forgets to mention.
This is also where ReAct-style reasoning (Yao et al. 2023) earned its keep for us. Letting the agent take an explicit "look at the schema first, then think, then act" turn made traces dramatically easier to debug. When something went wrong, we could read the agent's working notes and see the exact moment a wrong table was picked. The trace was the product as much as the answer was.
What Spider and BIRD do not measure
The metric that mattered to us did not appear in any leaderboard. It was something like: of the queries the agent chose to answer, how many survived a senior analyst's second look without correction? That is a harder bar than execution accuracy. It punishes confidence as much as it punishes error, and it rewards the agent for staying quiet when it should. We still report execution accuracy because it travels well in slide decks. But internally we track the survival rate, and we treat the abstain rate as a feature, not a bug.
There were evenings when I sat with traces open on one screen and PostgreSQL metadata on the other, wondering whether we were building intelligence or just packaging caution in a friendlier wrapper. Maybe the answer is both. Maybe most useful systems are less like prodigies and more like rituals that help humans make fewer avoidable mistakes. The users taught me that part. Analysts do not care whether the model used chain-of-thought shaped scaffolding under the hood. They care whether the answer survives a second look. They care whether the system embarrasses them in front of a manager. They care whether trust accumulates or leaks. That is a harder bar than fluency. I still like the magic of typing a question and watching data answer back. I just no longer confuse magic with reliability. An agent becomes useful the day it learns to say less.
References
- [1]Yu et al. (2018). Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task · EMNLP 2018The benchmark that taught the field that cross-domain generalization, not query length, is the real test.
- [2]Li et al. (2023). Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs · NeurIPS 2023BIRD pushed evaluation toward dirty schemas and external knowledge — closer to what production actually feels like.
- [3]Pourreza, Rafiei (2023). DIN-SQL: Decomposed In-Context Learning of Text-to-SQL with Self-Correction · NeurIPS 2023Decomposition plus a self-correction pass — the recipe behind a lot of agent loops since.
- [4]Gao et al. (2023). Text-to-SQL Empowered by Large Language Models: A Benchmark Evaluation · VLDB 2024A careful comparison of prompting strategies — useful when picking what to actually ship.
- [5]Yao et al. (2023). ReAct: Synergizing Reasoning and Acting in Language Models · ICLR 2023The reasoning-and-acting interleave that made our traces actually debuggable.
- [6]Wang et al. (2022). Self-Consistency Improves Chain of Thought Reasoning in Language Models · ICLR 2023Sample more, vote on the survivor — the cheapest reliability upgrade in the toolbox.
- [7]Anthropic (2024). Building Effective Agents · Anthropic Engineering BlogA short, opinionated essay; the line that stuck with me is that the best agents are the simplest ones.
- [8]Shinn et al. (2024). Reflexion: Language Agents with Verbal Reinforcement Learning · NeurIPS 2023Write down why the last attempt was wrong, then try again — surprisingly powerful as a critique loop.