Why Chatting With Your ERP Fails
October 6th, 2026 | by Lane Nelson
Part 4 of 10 – natural-language reporting, and what to build instead
The demo is genuinely impressive, and that is the problem.
A vendor is in the room at Spiese Fluid Power, showing off "chat with your ERP." Someone from the leadership team types a question in plain English – what were our sales to Kohler last quarter? – and two seconds later a number appears. No report to run, no analyst to wait on, no cryptic green screen. Just the answer, in a sentence. The room is delighted. Somebody says the quiet part out loud: why have we been paying for reports all these years?
Ellen, who has been controller here for fourteen years, is not delighted. She is doing the math in her head, and the number is wrong. Not wildly wrong – that would be easy. It is wrong by about eight percent, in the plausible direction, and she is the only person in the room who knows it.
That gap – between how good this looks and how quietly it fails – is the whole subject of this post. Natural-language reporting is the single most requested AI feature in every ERP shop, and pointed at your raw database, it is the one most likely to hand a confident wrong number to someone who will act on it. This is not an argument against it. It is an argument about how to build it so Ellen can stop being the last line of defense.
Why it demos beautifully and dies in production
The demo works because the vendor picked the question. Sales last quarter is clean: one table, one date, one sum. The model writes the SQL, the SQL runs, the number is right, everyone claps.
Production is not a curated question. Production is everyone in the building, asking whatever they actually want to know, in whatever words occur to them, against a database nobody designed for this. And there the accuracy falls off a cliff – quietly.
The evidence is now specific. When dbt Labs benchmarked language models writing SQL against raw tables in 2026, the best models got 64.5% of questions right. That is roughly double what they managed three years earlier, and it is the most dangerous kind of improvement – good enough to be trusted, not good enough to be trustworthy. More than a third of answers were wrong. And they were wrong the way Ellen's number was wrong: not with an error message, but with a plausible figure.
That is the property that matters. A report that crashes tells you it failed. A chatbot that hands you 41% gross margin when the real number is 34% tells you nothing – it looks exactly like the times it was right. You find out when the board asks why the quarter came in soft, or you never find out at all.
Benn Stancil put the general version of this plainly a couple of years back: language models should not be the thing writing your SQL.
I have given this demo before – or its grandparent. Starting in 1986, at Pansophic Systems, I demonstrated Easytrieve Plus – Natural Language: a still-pre-release option that took a plain-English question, answered it, and generated the Easytrieve Plus code behind it. Audiences loved it for the same reason the room at Spiese did – though it took a good deal of preparation to make each demo look that effortless. The language was genuinely hard back then – those early parsers were nothing like today's models – but that was not the real problem. Even a plain, unambiguous question could come back wrong, because the trouble lived underneath the words, in the layer of meaning between the question and the schema. The product understood that and made a serious attempt at the layer; it simply wasn't robust enough. Forty years on, the language is nearly solved, which only sharpens the lesson: what decides whether this works has never been the sentence. It is the semantic layer underneath – and it is still the part you cannot skip.
Your schema is the worst possible thing to point it at
Text-to-SQL is hardest exactly where ERP schemas live. The dbt benchmark used tidy analytical tables and still missed a third. A typical mid-sized manufacturer's ERP is not tidy. It is twenty-two years old, has a decade of customizations nobody documented, and a support staff trying to keep the lights on with no training in data engineering. And it carries the traps that make a model guess wrong:
The same word meaning five things. "Sales" is booked, shipped, invoiced, recognized, or net-of-returns depending on who's asking. The model picks one – usually the wrong one for the question – and never mentions there was a choice.
Status codes with tribal meaning. That third status code on manifold work orders – the one from the last post that only Dave knew the reason for – is invisible to the model. It will happily sum across a status that should have been excluded.
Field names a human can't read. Columns are six or ten cryptic characters – ODTOT, CSHFL3, WHSLOC – with no declared meaning. The model guesses from the name, and the name lies.
Ellen navigates all of this without thinking about it, because she has spent fourteen years learning which field is real. The model has spent two seconds, and it cannot tell you which of those traps it stepped in – because it doesn't know it stepped in one.
Where the boundary sits
Here is the mistake, in the language of this series: letting the inference layer reach directly into the system of record.
Every other post in this series puts the model on messy, linguistic input – a PDF, a help-desk ticket, a question about why. Those are its home turf. Chat-with-your-ERP does the opposite: it points the model at the most structured, most consequential, least forgiving thing you own, and asks it to produce numbers – the one output that has to be exactly right, and the one thing Post 1 warned never to ask a language model for.
The fix is not to abandon natural-language reporting. It is to stop letting the model touch the raw schema, and put something in between: a semantic layer. A thin, governed definition of your real metrics – what "sales" means, which date counts, which statuses are in, how margin is computed – with the joins and the exclusions and the business rules baked in, once, by someone who knows. The model no longer writes SQL against ODTOT. It picks from a short menu of defined metrics and dimensions, and the layer generates the query.
The same dbt benchmark measured this. Pointed at a semantic layer, the same models answered 100% of the questions the layer covered correctly – and when a question fell outside the layer, they said so, instead of guessing. Same models, same questions. The difference was entirely that someone had defined the metrics first.
What the AI does – and what it must never do
Does: turns the user's plain-English question into a selection against defined metrics and dimensions – sales, by customer, last quarter – and lets the semantic layer produce the arithmetic. It handles the language, which is what it's good at. The layer handles the numbers, which is what it's for.
Must never do:
Never write SQL against the raw tables. That is the whole failure mode – the improvised query is exactly where the wrong number comes from.
Never invent a metric. If nobody has defined "margin," the model does not get to decide whether it includes freight. It declines.
Never answer outside the layer's coverage. A question the semantic layer can't express is a question the model refuses – visibly – rather than guessing. The refusal is a feature; it is the thing raw text-to-SQL cannot do.
Never present a number without its definition. Every answer can say which "sales," which date, which exclusions – so Ellen, or anyone, can check it against what they meant.
The plumbing
The instinct is to buy "chat with your ERP" and point it at the database. That is precisely backwards.
Build the semantic layer first, and build it small. You do not need to model twenty-two years of schema. You need to define the ten or fifteen metrics people actually ask for – sales, margin, bookings, on-hand, past-due, DSO – each one nailed down: the real fields, the right joins, the status exclusions, the date that counts. That definition is the asset. It is also the hard part, and it is not an AI problem – it is Ellen and the data in a room, settling what the numbers mean, probably for the first time in writing. This approach assumes only what mid-sized organizations have: management that knows what numbers they need, and people like Ellen who know where to find them.
I learned how hard that room is long before AI existed. In 1988 I was hired onto the staff of AIM Rent-a-Car, a Budget Rent-a-Car franchisee, in part to fix a problem the president already knew he had: the company had five different answers to one simple question – what is our average revenue per rental day? Five systems, five numbers, all labeled the same thing, none agreeing, and no one able to say which was right.
Counting the rental days was the easy part, and even that had arguments – does a car rented twice on the same day count as two rental days, or one? The revenue side was worse. Was revenue the daily rate times the days on the agreement? Did it include the options – the insurance, drop fees, or car-phone rentals? Gross, or net of refunds and damages? And then the big one: the accounting version, where revenue is aggregated into monthly buckets to reflect earnings fairly with the benefit of hindsight, and different from the numbers reported as they happen. Every one was somebody's honest definition of "revenue," quietly baked into a different report.
Sorting it out took weeks, and almost none of the work was technical. That was 1988 – no AI within a mile of it. This failure is not an AI failure; it is decades old. AI just gets you the wrong one of the five faster, and says it with more confidence.
Only then does natural language go on top, pointed exclusively at that layer. The chat interface is the easy last mile. The governed definitions underneath are the product, exactly as the human-review step was the product in the order-intake post. And there is a quiet bonus: once "sales" is defined in one place, the five departments that compute it five ways finally have one answer – which may be worth more than the chatbot.
Failure modes and guardrails
The confident wrong number is the whole reason for the post. Guardrail: the model cannot reach a number except through a defined metric, so there is no path to a plausible improvisation. If it can't map your question to the layer, it says so.
Coverage creep. Users will ask for metrics the layer doesn't have, and the pressure will be to let the model "just try" for the uncovered ones. Don't. The moment you allow improvised SQL for the long tail, you are back to 64.5% and silent failure – now hidden behind a layer people have learned to trust, which is worse.
The false sense of governance. A semantic layer that is wrong is still wrong, confidently and at scale. The metrics need the same review any published number gets, and a way to correct a definition once and have every answer follow.
How you'd know it worked
Not user delight – that shows up in the demo and tells you nothing. Two harder measures.
First, accuracy on a fixed question set. Write down the fifty questions your people actually ask, work out the true answers with Ellen once, and test against them on every change. The target is not "high." For numbers that go in a board deck, the target is 100% within the layer's coverage – which is achievable, because that is what defining the metrics buys you.
Second, and stranger: the refusal rate on out-of-scope questions. A system that never says "I can't answer that" is not being helpful – it is guessing, and you should trust it less, not more. The willingness to decline is the single clearest signal that the boundary is holding.
Monday morning
Do not buy chat-with-your-ERP and point it at your database. You now know exactly how that ends.
Instead, spend an afternoon on a list. What are the ten questions your leadership asks about the business most often? Sales by customer, margin by product line, what's past due, what's on hand, how are we tracking to plan. Write them down.
Then take the top three to your Ellen and settle what they actually mean – which fields, which dates, which statuses in and out. Define those three, precisely, in one place. You will spend most of the time arguing about definitions, and that argument is the most valuable part; it is work you needed to do anyway and have been avoiding.
That is the whole first project: three metrics, defined and agreed. No chat interface yet. Once the definitions are solid, natural language on top is a weekend. Get that order wrong – interface first, definitions never – and you have built a very expensive way to be confidently wrong, with Ellen as the only thing standing between it and the board.
Define the numbers first. Let the model do the language.
Spiese Fluid Power is fictional – a composite built from published industry benchmarks and anonymized operating ratios. Its full operating profile, with sources, is published separately.