Most AI Context Layers Are an Inch Off. That's a Mile Wrong.

Every context layer is a customs house. Most stand an inch from the border that matters, where runtime context lives.

7 min read

In almost every conversation I have with a CIO, the concern is the same: an agent gives a confident, wrong answer and someone acts on it. A wrong answer in a demo is a curiosity; in production it's a bad forecast, a missed fraud flag, or a customer told something that isn't true. If you give a model a raw schema, it has to guess what your company means, and in the enterprise a confident guess is worse than no answer at all. So I agree with where the industry is heading: agents need a context layer that carries the company's own definitions, relationships and rules. Semantic models, metric stores and ontologies are all being rebuilt for exactly this. The instinct is right. But most of that work is aimed at the easier half of the problem, and it's being built in the wrong place.

Most AI Context Layers Are an Inch Off. That's a Mile Wrong.

The context layer is a customs house

Enterprise data has never been more reachable. The harder problem is knowing what it means. Ask an agent for last quarter's revenue and it will give you a number. Whether it's the number your CFO would sign is another question, because revenue can mean bookings, ARR, recognized revenue or cash collected depending on who is asking. The same customer may appear differently across systems, and a metric that looks obvious in a dashboard may carry exceptions that live only in finance, sales or operations. Give an agent all of the data without that context and it will produce an answer that is well-formed, confident and wrong.

This is the last inch between data and intelligence, and the context layer is the customs house that guards it. It supplies the definitions, relationships, rules and permissions that tell an agent how the business actually works. The better that context is, the less the system has to infer for itself and the fewer plausible-but-wrong paths it can take. 

Design-time context is the easy part. Runtime context is not.

Context comes in two kinds, and only one of them is hard. Design-time context is everything that can be defined in advance: schema descriptions, approved joins, metric definitions, business rules, permissions and golden queries. A team can curate it, test it and teach it to the system before the first question is asked. Most serious platforms now support it, ours included. It’s the customs code, and writing it down is a largely solved problem.

Runtime context is the cargo. It’s created after the system starts moving: the reservation made two seconds ago, the payment that just cleared, the fraud signal that just arrived, the inventory adjustment another agent just wrote. From that point on, the next correct action depends on what the last one changed. None of it can be specified in advance, because it’s the changing state of the business itself. That makes recently written and ingested data part of context, not a freshness concern, and it is the most overlooked part of context in enterprises today.

That distinction decides where the customs house belongs. Most context layers are being built downstream, on analytical copies fed by change data capture. They carry design-time context perfectly well. What they cannot carry is runtime context: the state created by the transactions, events and other agents that have acted since the copy was made.

When agents read the wrong state

A read from a replica returns the same confident result whether the relevant write landed three seconds ago or seven, and nothing in the result tells the agent which case it's in. Four things make this worse than a simple delay: the lag moves, agents cannot see their own writes, the copy can show states that never existed, and two engines can disagree about the same row.

Replication lag is usually a few seconds and occasionally minutes, and the spikes arrive with bursts of writes: a flash sale, a fraud wave, a market move. The copy falls furthest behind exactly when the data is most active and correctness matters most. The agent never sees the lag. It sees a number.

Agents also can't reliably see their own writes. Say one reserves the last unit, then checks inventory through the replica two seconds later. The reservation has not arrived yet, so the unit still looks available, and the agent reserves it again. Across a fleet it gets worse, because a second agent sees the same phantom unit and books it too. That is not a stale answer. It is a double booking with a real cost, and it compounds with every agent acting at once.

Sometimes the copy shows a state that never existed at all. CDC applies changes as a stream, so at any given instant the replica can hold the debit from a transaction without the matching credit. That read is not a delayed snapshot of the source. It is a snapshot of a world that never was. Some pipelines guarantee transactionally consistent reads and some do not, and most teams running one have never checked which kind they have.

Even perfect replication would not solve the problem, because two engines can still give two answers. With zero lag, an agent querying the replica and an application querying the source can disagree about the same row at the same moment. Consistency across agents comes from one engine working over one copy of the data, not from replicating faster.

Each of these failures comes down to an agent reasoning over one version of the business and acting on another, and nothing tells it the two have drifted apart. That was tolerable while AI only answered questions. When an agent that answers questions is wrong, the damage lands in a slide deck. When an agent that writes back is wrong, it lands in production. Correctness is not something a better model can supply. It is a property of where context lives.

The customs house belongs where the writes happen

Context layers ended up downstream for a reason. Operational databases were built to record transactions, not to run analytics, search and vector retrieval across them, so every team copied the data somewhere that could. That tradeoff is what created the gap, and it's the constraint we set out to remove.

Putting the customs house where the business transacts means keeping definitions, relationships, golden queries and permissions on the same live data the applications write to, and letting transactions, analytics, full-text search and vector search run against it together. One engine, one copy, one version of the row. That is the bet we have made at SingleStore: agents reason over the same current rows the business is still writing, and their actions flow back into that same state. It is also why Domains in Aura Analyst carry semantic context and database privileges together, so an agent's definitions and its permissions are enforced on the same live rows.

Keeping everything in one engine also removes most of the border crossings. Watch an agent work today and much of what looks like reasoning is shuttling between systems: pulling schemas, querying the warehouse, checking the operational database to see whether the answer still holds, and reconciling identifiers that almost match. The duty on every crossing is paid in tokens, and none of it makes the answer more correct.

Any data leader can test their own stack with three questions. Can your agent read its own writes? Can two agents asking the same question at the same moment get different answers? Are your business definitions and your permissions enforced where the transactions happen, or on a copy? If the answers are no, yes and on a copy, the context layer is a reporting tool with a new name.

Nearly everyone is building a customs house. The question worth arguing about is which border it stands on. Build it downstream and it inspects cargo that has already shipped, and the agent reading from it ends up where dashboards are today: useful for understanding what happened, insufficient when something has to be decided. Build it where the business transacts and every agent crossing over has to declare what it means, what it may touch and whether it still describes the world as it exists now. That is the last inch and it’s where the answers go right or wrong.


Share