Why do AI agents give different answers on the same data?

Same model, same question: an agent querying sources directly reaches 38% accuracy over 234 tool calls, versus 91% accuracy in one SQL query on a governed data layer.

Because a connection protocol like MCP tells an agent how to reach your data, not what it means. Without a governed semantic layer, an agent reinvents a definition on each run, so the same question returns different numbers. The fix is a Raw, Trusted, Service data layer plus Context Engineering.

Why do AI agents give inconsistent answers on the same data?

Because most agents are connected to data without being told what it means. A protocol like the Model Context Protocol (MCP) standardizes how an agent reaches a source, but it carries no definition of what an "active customer" is, which date counts as revenue, or whether two records with different names are the same company. When that context is missing, the model does the only thing it can: it fills the gap with a plausible assumption. Run the same question tomorrow and the assumption changes, so the number changes with it.

The failure is not a weak model. It is business logic being improvised inside a context window instead of being defined once, in the data.

Same model, same question, two architectures. Querying sources directly, an agent answered correctly 38% of the time, across 234 tool calls, in about 15 minutes. On a governed data layer, the same agent reached 91% accuracy with a single SQL query in 19 seconds. (Nekt benchmark)

Does MCP solve this?

No. MCP solves connection, not meaning, and those are different problems. It standardizes how an agent discovers and calls a tool, which is a real advance: before it, every integration was bespoke, and after it, integration became a protocol. But knowing how to talk to a system is not the same as knowing what its data means. A source can hand an agent the raw value of a field. It cannot tell the agent that the field is stored in millionths, or that a deal marked "won" two years ago should not be counted as revenue today. That knowledge is context, and context does not travel in the protocol. Someone has to model it.

What goes wrong when the semantic layer is missing?

The agent turns into an improvised database made of text, and it fails in four ways that never show up in the demo. Ask it "how much did each new customer cost last quarter, by channel," and with nothing underneath, it has to page through the CRM, pull contacts (because the channel lives on the contact, not the deal), pull ad spend from several platforms in several formats, pull billing to know who actually paid, and then join all of it by reading records inside the context window. That last step is where it breaks: it mismatches time zones, counts duplicates, overflows the window, and loses half of what it already read.

  • It invents the definition. Asked "how many active customers," with no rule it makes one up: live subscriptions on Monday, won deals on Wednesday. Two defensible numbers, both different. The reader loses trust and returns to the spreadsheet.
  • It joins in the wrong place. Combining tables is database work: keys, indexes, computation. Doing it by reading inside the context is expensive and fragile, and it is the main source of the token blowup.
  • Scope leaks. Without per-agent permissions at the data layer, access control becomes a line in a prompt. A prompt is a polite request, not a boundary.
  • Cost grows with volume, not with the question. The dataset doubles, the token bill doubles, and the answer is no better. This is where projects die on budget rather than on engineering.

The scale of that risk is now on the analyst radar. Gartner projects that over 40% of agentic AI projects will be canceled by the end of 2027, citing rising costs, unclear business value, and inadequate controls. Rising cost and unclear value are precisely what an ungoverned data path produces.

What data layer fixes it?

A three-stage layer, where each stage has one job, so the agent reads meaning instead of rebuilding it. The stages:

  • Raw keeps what the source sent, unchanged, so any number can be audited back to its origin.
  • Trusted holds the cleaning that no one sees and everyone depends on: deduplication, correct types, correct time zones. One real example: when two records are merged in a CRM, the absorbed record often stays in Raw because the connector never captures the deletion, so counting from Raw counts the same customer twice. The rule that filters it lives in Trusted, written once and applied to every agent and every report.
  • Service is the data modeled in the language of the business: not "deals table" but revenue_by_channel or active_customers, one table per question the company actually asks. This is where the agent reads.

The effect is that intelligence moves out of the prompt and into the data model. The agent stops being the place where business rules live and goes back to what it is good at: understanding the question in plain language and choosing the right table. One query, 19 seconds.

What is Context Engineering, and why do agents need it?

Context Engineering is the practice of writing a column's business meaning next to the data itself, so an agent reads the rule instead of guessing it. A person who sees a strangely named column asks a colleague. An agent does not ask, it deduces, and it deduces with confidence. So every column that matters carries an annotation. Real cases that break agents without it:

  • Ad cost stored in millionths, where one currency unit is 1,000,000. Without the note, the agent reports a cost per acquisition a million times too high, and explains the crisis convincingly.
  • Email type, separating corporate from personal. Corporate converts at 15 to 20%, personal near 1%. An agent that blends them concludes that a healthy channel is broken.
  • Historical wins. A customer who churned still has a win recorded in that month; if they return in another year, there are two. Without the rule, the agent reports "never a customer" about someone who is paying right now.
  • Which date is the date. Created, closed, first charged: three dates, three answers to "how much did we sell in June," and one of them becomes a salesperson's commission.

None of this is data. All of it is context about the data, and it is exactly what a connection protocol does not carry. The discipline is old, since giving data meaning has always been data work. What changed is the consumer: it used to be an analyst who asked when unsure, and now it is an agent that answers even when it does not know.

In practice, this context can be served automatically. When an agent queries Nekt over MCP, it receives more than the table. It receives the definition alongside it: how the company measures MRR, what counts as an active customer, which date becomes revenue. The business rule arrives embedded in the answer instead of being left to a well written guess.

How do you build a data layer for AI agents?

Start from the questions, not the tools. The order that works:

  1. Write the ten questions the business actually asks every week. That, and nothing more, is the scope of the Service layer.
  2. For each question, list the sources and the key that joins them. If no key appears in all of them, that is the first problem, and it predates any AI.
  3. Bring everything raw into Raw, with no cleaning, so it can be audited and reprocessed.
  4. Write the tedious rules into Trusted: deduplication, time zones, types, merged records, currency units. Once, in code, versioned, never in a prompt.
  5. Model Service in the language of the business, one table per question.
  6. Annotate context on every column a human would have to ask about.
  7. Only then connect MCP, with per-agent permissions and minimum scope. MCP is the last step, not the first, and starting there is why so many agent projects stall.

The conclusion is not "MCP or a semantic layer." It is three things together: structured data underneath, MCP to connect, and Context Engineering to give meaning. Remove any one and the agent returns to guessing. The model is the same at 38% and at 91%. What changes is what sits beneath it.

  • Inconsistent agent answers come from missing business context, not a weak model: without a defined semantic layer, an agent reinvents definitions on each run.
  • MCP standardizes connection, not meaning. It is necessary but not sufficient on its own.
  • A Raw, Trusted, Service layer moves business logic out of the prompt and into the data, turning 234 tool calls into a single query.
  • Context Engineering, writing each column's meaning next to the data, lets an agent read the rule instead of guessing it.
  • Build from the ten questions the business asks weekly, and connect MCP last, with per-agent, minimum-scope permissions.

More insights