What is a data contract, and how do you enforce it across systems?

A CRM record owned by three teams (Sales, Finance, CS) becomes one rule checked daily: payment, won deal and post-sale mirror joined by one key.

A data contract is an explicit agreement about what a piece of data means, who produces it, and the rules it must always obey, written before the code rather than discovered after it. Enforcing one across systems means unifying the sources into a governed layer by a shared key, then having an agent check and repair every record against the rule, part language-and-judgment handled by a model, part arithmetic-and-verification handled by deterministic code. This guide covers what a contract specifies, how to enforce it on a CRM, and where the AI should stop and code should take over.

What is a data contract?

A data contract is an explicit agreement about what a piece of data means, who produces it, and the rules it must always obey, written before the code instead of discovered after it. For a CRM, the contract is a statement that must always hold for a paying customer, and it is checked by machine every day. Writing the rule first turns a project inside out: instead of programming "do this, then that," you program "how do I detect and repair the moment someone breaks the rule." And because the rule lives in the repository rather than in one person's head, it keeps holding when that person is away.

A data contract is a rule that must always be true, written before the code and verified by machine every day. The job stops being automation and becomes detection and repair: catch the record that violates the rule, and fix it.

The idea is old in data engineering, but it matters more now that an agent, not an analyst, reads the result and acts on it. A contract is the difference between a CRM that can be trusted to drive a forecast and one that quietly disagrees with itself. The rest of this guide uses the CRM as the worked example, because it is the asset where the contract pays off first.

Why does CRM data rot?

Because it has three owners and no one responsible for consistency. A single customer is touched by three teams, each in its own system: Sales closes the deal, Finance bills the payment, and Customer Success delivers the service. Each edits its own system without looking at the others, so the CRM either starves (the deal never moves to won, the post-sale record is never created) or overeats (the same customer duplicated, with dates that disagree). The failure is one sale at a time, always slightly out of conformance.

This is not a people problem; it is a data problem. Joining the four systems is not the job of Sales, or Finance, or CS, so nobody does it, and the record drifts. And the CRM is the worst place for this to happen, because after the product it is the most important asset in the company: forecast, attribution, quota and commission all come from it. When it loses trust, the damage compounds in a specific way. People stop checking commission against the CRM and start auditing it straight from Finance, which means the most important tool in the company gets bypassed on the number that matters most. The pattern of every tool reporting its own version is the subject of why AI agents give different answers on the same data, and the same CRM, trusting an inferred field instead of the underlying data, is what under-reports your paid leads.

What does the contract actually specify?

It specifies the sides that must exist together and the key that joins them. In the worked example, a paying customer must have three things at once: a payment, a deal marked won, and a mirror record in post-sale. All three, always, joined by one key, the customer's workspace (in many businesses this key is the tax ID). Without that key, the payment, the deal and the post-sale record are three strangers that never meet.

Answering whether the three sides hold means reading across four systems and twelve tables, joined by that single key.

SystemWhat it holdsIts side of the contract
CRMDeals and contactsWho closed, and is the deal in won
BillingCharges, customers, subscriptionsIs there a real payment, and when
ProductOrganization, subscription, usageThe workspace key that joins everything
SupportConversations, including messaging appsEvidence of handoff and follow-up

A join is combining two tables by the key they share. It is easy with two. The contract requires it across twelve tables from four systems, every day, over the whole base, without the source APIs rate-limiting you and without each run returning a different number. That is why nobody does it by hand, and why a governed layer, not a spreadsheet, is the only place it holds. A verification that once took at least ten minutes per sale runs automatically.

How do you enforce a data contract?

With an agent that runs the rule over every record and repairs what it can, built as two agents that cover each other. One is live: it wakes when a sale is announced and runs the rule on that sale immediately. The other runs daily over the entire base, re-checks the rule on every customer, and produces a report of what broke today, what has been broken for days, and what was resolved since yesterday. The daily sweep is the net for anything the live agent missed, and it is what makes "no human owner" real, because it runs on a schedule whether or not anyone is watching.

In practice the agent turns the contract into a checklist of ten objective items run against each sale: is the deal in won, is the close date correct, is the source filled in, does the post-sale mirror exist, is the subscription matched to the workspace, and so on. Each item has one of three answers: yes, no, or unsure. What is clearly resolvable, the agent fixes. What is unsure, it does not guess; it asks the right person and does not drop the thread until it is closed. Underneath the checklist are five questions, each living in a different system, that together decide whether a customer truly exists:

  • Who closed, and who is the customer? The deal and its owner. Recognizing that a deal name, a legal name and a trade name are the same company is judgment, not arithmetic.
  • When did it close, and when did the money arrive? This date becomes commission, so getting the day wrong moves someone's pay. Timezone errors live here.
  • Where does the customer live? The workspace key that joins every system.
  • What, how much, and how do they pay? The subscription and the charge that actually cleared, which separates a paying customer from a deal marked won on optimism.
  • Why is there no follow-up yet? If a payment and a won deal exist but the post-sale mirror does not, the customer is in limbo: paying, and invisible to the team meant to serve them.

Where should the AI stop and deterministic code take over?

Draw the boundary by capability: the model owns language and judgment, the code owns arithmetic and verification. Put everything on the model and it stamps a wrong date, forgets a case, and hands back a number it "decided," which you then trust. Put everything in fixed rules and the real world has more variation than your list, so it becomes an endless game of patching exceptions. The durable design splits the two cleanly.

The model ownsThe code owns
Reading intent and deciding if a case is clearChecking that the three sides of the contract exist
Seeing that two similar names are the same companyDate math with the correct timezone
Writing the right message to the right personReading a field, dismissing a false alarm

The mechanism that keeps the two apart is a structured seal. The model understands a messy case once and writes down a structured mark, a clean verdict; the deterministic code only reads that mark and applies it. The code never parses natural language, because parsing intent in code is exactly the brittle work the model exists to do. There is a useful rule of thumb: when the system makes a mistake, the intelligence is almost always on the wrong side of the boundary. This is the same principle as keeping business meaning in the data rather than the prompt, covered in context engineering for AI agents, and it is why the code addresses pipeline stages by ID, never by name: the name changes, the ID does not.

What makes the enforcement durable, not a script that rots?

Five things that either exist or do not. Running an agent alone in the cloud is not the same as running it in a terminal, where memory, orchestration and context come for free; in the cloud the model arrives raw and that scaffolding is built by hand. The part that looks like AI is the easy part. What holds it up underneath is the work, and it comes down to five requirements.

  • Versioned in Git, so anyone on the team can open it, not kept on one laptop.
  • Running on the company account, not on a personal AI session.
  • Deployable by any engineer, so a version tag ships it, even with the author away.
  • Scheduled, so it runs every day, including during holidays.
  • Recognized as a real service by the engineering team, not treated as someone's personal script.

Without these it is a session, not a service. A service has collective ownership, and that is what makes "it runs without me" true rather than aspirational.

How does Nekt do this?

Nekt is a data platform that gives AI agents governed context, and enforcing a data contract is a direct example of the three parts it covers.

  • Connect. The CRM, the product database, billing and support land in one place through ready-made connectors, so a support message that would otherwise die on someone's phone enters as just another source.
  • Govern. The sources are unified by the shared key, the twelve tables joined once, and the contract itself is modeled in the layer with business meaning annotated per column, so the rule is written in the data rather than re-described in each query.
  • Serve. The agent reads the result over an MCP server and checks every record against the contract, with the model handling judgment and deterministic code handling verification. The same layer that reads the data is the one that lets the agent repair it.

The payoff is a record that stops decaying in the dark: the most important asset after the product, the one every decision is built on, is checked against its contract every day. See how Nekt works, or create a free account and connect your first source.

Frequently asked questions

What is a data contract in simple terms?

It is a written rule about what a piece of data must always look like, agreed before the code is built and checked automatically. For a CRM, a contract might say that every paying customer has a payment, a won deal and a post-sale record, all joined by one key. If any side is missing, the contract is violated and the record needs repair.

Data contract vs data quality: what is the difference?

Data quality is the general state of the data being correct, complete and consistent. A data contract is the specific, written rule that defines what correct means for a given dataset, and the mechanism that enforces it. Quality is the goal; the contract is the agreement and the check that gets you there, so it does not depend on anyone remembering the rule.

Should an AI agent fix CRM data automatically?

For the clear cases, yes; for the ambiguous ones, no. The durable design lets the agent repair what is objectively resolvable and escalate what requires judgment, rather than guess. The AI decides whether a case is clear and writes a structured verdict; deterministic code applies the fix. Letting the model invent values it is unsure about is how a repair tool becomes a new source of errors.

  • A data contract is an explicit, written rule about what data means and must always obey, defined before the code; enforcement becomes detecting and repairing violations, not building automation.
  • CRM data rots because it has three owners (Sales, Finance, CS) in separate systems and no one responsible for consistency; the fix is to treat the CRM as data and unify the sources by a shared key.
  • The worked contract is three sides joined by one key (payment, won deal, post-sale mirror), which requires joining twelve tables across four systems every day, over the whole base.
  • Draw the boundary by capability: the model owns language and judgment and writes a structured verdict once; deterministic code only reads that verdict and handles arithmetic and verification, addressing stages by ID, never by name.
  • Durable enforcement needs five things: versioned in Git, run on the company account, deployable by any engineer, scheduled, and recognized as a real service, or it is a session, not a service.

More insights