Governed Datasets as the Agent Data Contract

Agents acting on data need the implicit knowledge humans carry, baked into the data itself.

Contributing Editor · · 11 min read
Cover illustration for “Governed Datasets as the Agent Data Contract”
Agentic Data Infra · October 6, 2026 · 11 min read · 2,533 words

A human analyst looking at a revenue number that seems too high will pause, check a filter, maybe ask a colleague before putting it in a slide. An agent given the same number acts on it, confidently and immediately, because nothing in its design tells it to doubt a query result. That gap in behavior is the reason the data infrastructure built for analytics over the last two decades cannot simply be handed to agents and trusted to work.

Human analysts carry context that was never written down anywhere. They know which tables are half-populated and best avoided, what "revenue" actually means at their specific company once returns and timing adjustments are factored in, and that a Tuesday in early January might fall on a holiday that throws off a week-over-week comparison. None of that lives in the database. It lives in the analyst's head, applied silently, query after query, as a kind of unpaid labor that the data layer never had to account for. A 2026 analysis on martinfowler.com, "Making Your Data Ready for Agentic AI," names this as the central design problem: once an agent replaces the analyst, every bit of that implicit work, meaning, sanity checks, judgment, has to move into the data itself, because there is no longer a person there to supply it for free.

The scale and speed at which agents operate make the problem worse, not just different. A human analyst runs one query, looks at it, thinks, and maybe runs another an hour later. An agent can call a metadata tool, pull a metric, check whether it's fresh, trace its lineage, run a follow-up query, and generate a final answer, all within a single automated pass that takes seconds. Every flaw sitting in the underlying data gets carried through that entire chain rather than caught partway, because nothing in the chain pauses to ask whether something looks wrong.

The stakes of that difference appear directly in production. Ampcome's governance analysis points out that a stale or inconsistent metric definition used to be a reporting annoyance, something a human might notice and shrug off. Once an agent is acting on that same metric, it becomes an operational error: a wrong alert sent to the wrong team, an escalation that should never have fired. OvalEdge's 2026 agentic governance guide makes a related, structural point: governance systems built around periodic human review, quarterly audits, scheduled spot checks, simply run at the wrong speed once agents are making decisions in real time.

Requirements of agent-ready data beyond AI-ready data

Diagram: AI-Ready vs. Agent-Ready Data: Five Key Differences. Visualizes: Contrast two paradigms side by side across five dimensions drawn directly from the article: (1) Consumer — ML engineer vs.

Agent-ready data is a different category of thing from AI-ready data entirely, built to satisfy a different consumer with different failure modes. Nexla's 2026 guide to data for AI agents lays the contrast out directly. AI-ready data was built for an ML engineer working with tables and features, governed before it ever reached the warehouse, refreshed on a batch schedule, and prone to failing quietly in the form of a bad prediction. Agent-ready data serves an autonomous agent acting at runtime, arrives as discoverable data products instead of static tables, enforces governance at the moment the agent makes its call, carries a freshness SLA as an explicit term instead of a batch assumption, and fails loudly, so an agent working with properly built data refuses, escalates, or retries when something is wrong.

That shift demands five things that human analysts never needed the data to provide, because analysts supplied them on their own. The first is discovery: a registry of what data exists, what it means, and what any given agent is permitted to touch, without which every new agent integration becomes a one-off project built from scratch. The second is semantic context, and the example is almost mundane in its clarity: a column named acct_st tells an agent nothing, while a data product labeled "Customer Account Status, verified daily, owned by Finance" can actually be called and trusted. The third is governance that survives the full pipeline, not just the warehouse, since agents often reach data through tool calls and protocols rather than the original, carefully permissioned query path. The fourth is a freshness guarantee: agents acting on stale data cause more damage than agents that simply decline to answer, so agent-ready data needs a known time-to-live, with the pipeline itself telling the agent when an answer has expired. The fifth is a runtime interface the agent can actually call, not a capability documented somewhere in a specification file that nothing in the agent's workflow ever reads.

None of these five properties can be bolted onto a data warehouse piecemeal and expected to hold up. They have to arrive bundled together, as a single artifact, before an agent ever lays eyes on the data. That bundling is what a governed dataset, treated as a contract between the data layer and the agents that consume it, is built to do.

What a governed dataset as a contract must contain

Diagram: Five Clauses of a Governed Dataset Contract. Visualizes: Show the five clauses that must all be present for a governed dataset to function as a contract, drawn verbatim from the article: (1) Schema as law — enforced validation, quarantine…

A governed dataset works as a formal contract between the data layer and any agent that consumes it. Like any contract, it holds up only if every clause is present, and dropping even one clause doesn't weaken the agreement, it voids it.

The first clause is schema as law. Thoughtworks' agentic data readiness analysis describes data contracts as code: the schema is an enforced rule rather than a polite suggestion that occasionally gets honored, because an agent treats every value it receives as ground truth and has no instinct to question the shape of what it's handed. That requirement produces what's called the quarantine pattern: data that fails schema validation simply never reaches the layer an agent is allowed to query, closing off the possibility that malformed or mistyped data quietly reaches an automated decision.

The second clause covers business-context definitions, the semantic layer that turns a column of numbers into something an agent can reason about correctly. A metric gets defined once and compiled directly into the dataset, rather than reconstructed independently by every agent that happens to query it, which would produce as many versions of "revenue" as there are agents asking about it. Thoughtworks' analysis uses revenue as its working example: the dataset itself has to encode that the figure already has returns netted out and that the fiscal year begins in February, details a human analyst would have carried around without ever writing them down. Ampcome's governance analysis shows what happens when this clause is skipped: when "gross margin" means something slightly different across two subsidiaries, a human looking at a dashboard treats the discrepancy as a number worth a second look, while an agent acting on the same inconsistency fires a wrong purchase-price alert without hesitating.

The third clause is access control that travels with the data and does not stop at the database's front door. Role-based access and row-level security have to be properties of the dataset itself, because agents frequently reach data through tool calls and through protocols like MCP rather than through the original, carefully scoped query path a human analyst would have used. Ampcome's analysis frames the stakes directly: an agent inherits the reach of whatever credentials it holds, so the boundary on what it may read and write determines how bad its worst day can get, and that boundary belongs inside the dataset contract rather than hidden in a prompt instruction that an agent can be talked out of. Decube's agentic governance guide reinforces the same structural point: access scope and dataset curation govern what an agent may read, what it may write, and what it has to be able to prove it did afterward.

The fourth clause sets freshness guarantees on a per-consumer basis, because a single dataset can carry very different freshness requirements depending on who's asking. A pricing table updated once a night might be perfectly fine feeding a dashboard a human checks each morning, but a quoting agent pulling from that same table may need near-real-time updates to avoid quoting a stale price. The freshness SLA is a binding term of the contract rather than a metadata footnote attached for documentation purposes, and an agent that doesn't know it's reading yesterday's pricing table isn't making a judgment call, it's making a mistake it had no way to catch.

The fifth clause is an auditable lineage trail. Agentic lineage, a traceable record of what data an agent touched, when it touched it, and under which version of the contract it was operating, has to be built directly into the dataset layer. Reconstructing that trail after an incident, from logs scattered across a pipeline, is not the same guarantee as having it recorded at the moment the agent acted.

How Ad-Hoc Query Access Breaks the Contract

Handing an agent direct query access to a production database doesn't create a weaker version of the contract above. It eliminates the contract outright, because every one of the five clauses depends on a pre-modeled artifact existing before the agent ever queries anything, and a raw production table is precisely what has not been pre-modeled.

Each clause fails in its own specific way. The schema clause fails because an agent querying a live production table gets whatever the schema happens to be at that exact moment, with no validation layer sitting in front of it and no quarantine catching a column that quietly changed meaning since yesterday. The semantic context clause fails because raw tables carry no business definitions at all: an agent has to infer what "revenue" means from column names and sample values, and it will infer something slightly different every time it's asked, the same instability Ampcome's gross-margin example describes, now happening on every single query.

The access control clause fails in a way that has already been documented at scale. In Supabase and Postgres, auto-generated APIs expose an entire table the moment it's created through SQL or a migration, unless Row Level Security is explicitly turned on, and even then, enabling RLS without an accompanying policy locks out all access. Leaving RLS off, meanwhile, exposes the full table to anyone who can reach the API. Security researchers documented this failure pattern at scale as CVE-2025-48757, affecting more than 170 AI-generated applications that shipped without RLS policies in place. The freshness clause fails because a raw table carries no SLA at all: an agent querying it has no way of knowing whether the data was refreshed an hour ago or has sat stale for a week. And the lineage clause fails because a raw query against a production table leaves no governed audit trail behind it, meaning what the agent touched, why it touched it, and under which metric definition it was operating at the time becomes unrecoverable after the fact.

This isn't a hypothetical risk confined to poorly run teams. Supabase has said that the majority of new databases on its platform are now created by AI coding agents rather than by a human writing a schema by hand, which turns the RLS gap from an occasional misconfiguration into a structural exposure built into the platform at scale. Ampcome's analysis frames the consequence in terms of blast radius: the worst outcome of ad-hoc agent access is bounded entirely by whatever credentials the agent happens to hold, and without a governed dataset layer sitting in front of it, that boundary isn't set by policy at all, it's set by accident, by whatever the schema happens to expose on a given day.

The performance case points in the same direction as the governance case. An agent running analytical queries directly against a transactional Postgres instance competes for resources with live application traffic trying to serve real users. Pre-calculated datasets stored in governed Parquet format and queried through DuckDB run fast and cheap without touching the production system. Isolating analytical workloads from transactional ones is a basic architectural rule, and skipping it is among the most common failure patterns found in production AI systems. A governed dataset doesn't just make agent behavior safer in some general sense. It's the only architecture that makes the worst case knowable in advance, because the boundary on what can go wrong is written into the contract rather than left to whatever a raw schema happens to allow.

MCP as the delivery mechanism for the contract, not a substitute for it

The Model Context Protocol has become the standard way agents discover and call data, and that has led to a common misreading of what it actually solves. MCP moves the contract to the agent. It does not create the contract, and an agent connected through MCP to ungoverned tables will still return confidently wrong answers, just through a cleaner interface.

What MCP actually provides is a runtime surface: a way for agents to discover data products and call them, satisfying the "reachable via a tool call" requirement that agent-ready data needs. The July 2026 spec revision, dated 2026-07-28, brought a stateless protocol core along with an Extensions framework, Tasks, MCP Apps, hardened authorization, and a formal deprecation policy. Making the protocol stateless matters directly for the fanout situation described earlier: the kind of multi-step agent workflow that calls a metadata tool, retrieves a metric, checks freshness, and runs a query in one pass becomes horizontally scalable once the protocol itself doesn't need to hold session state between calls. The GitHub MCP Server has already moved to stateless operation and removed Redis session storage entirely, eliminating database writes on initialize and database reads on every subsequent call.

None of that infrastructure improvement touches the issue of what's sitting behind the protocol. MCP exposes exactly whatever it's pointed at. If what's behind it is a raw table with no business definitions, no freshness SLA, and no governed access policy, the agent receives all of that ungoverned, undocumented data through a protocol call that looks clean and structured on the surface. The semantic layer is what makes the answer an agent gets through MCP trustworthy. MCP is what makes that semantic layer reachable. Neither one substitutes for the other.

A well-built MCP analytics server reflects that division clearly in what it chooses to expose. It offers tools that describe approved datasets and state their contract terms up front. It offers tools that retrieve metric contracts, the compiled business definitions that make up the second clause of the governed dataset contract. It runs governed metric queries against pre-modeled datasets. It checks a table's freshness against the specific SLA set for that consumer, satisfying the fourth clause, and it exposes lineage for a given asset, satisfying the fifth. It also supports comparing current values against baselines and routing requests through workflow approval when a decision calls for one.

Security considerations reinforce why this division has to hold. MCP's own guidance describes "confused deputy" risks that arise in proxy servers and states that token passthrough is forbidden. A governed dataset layer, built with explicit access controls as one of its clauses, is precisely the architecture that contains that risk, because the agent's permissions are bounded by the terms of the dataset contract itself rather than by whatever the MCP server in front of it happens to allow through. The protocol can only carry a guarantee that the dataset behind it already provides.

Sources

  1. Architecture overview - Model Context Protocol
  2. The 2026-07-28 MCP Specification Release Candidate
  3. Scaling AI Agent Infrastructure with the MCP Stateless updates - Google Developers Blog

More in Agentic Data Infra