Why enterprise AI agents can behave perfectly and still reach the wrong conclusion
The agent gives a perfectly reasonable answer.
The model behaved as expected. The prompt was fine. The tool call worked.
And the result is still wrong.
I think this is where a lot of enterprise Agentic AI projects are going to get interesting.
Because once agents start working with operational data, many of the failures will have very little to do with the model itself.
They will look much more familiar to data people.
Wrong timing. Incomplete state. Conflicting systems. Missing history. Business meaning that exists in people’s heads but nowhere the machine can reliably find it.
The AI may simply be where the problem becomes visible.
Microsoft Research’s recent work on contextual integrity in enterprise agents had similar findings. Its CI-Work research found that increasing model capability alone did not remove information-handling problems and argued for more emphasis on context-centric architecture, rather than relying purely on better models. Microsoft
That feels increasingly important.
Everything can be correct and the answer can still be wrong
Imagine an engineering agent trying to determine whether a piece of equipment can return to service.
The maintenance system says the repair is complete.
The inspection system records a successful test.
The parts system shows the replacement component installed.
The operational planning system still shows the asset unavailable.
Nothing there is necessarily incorrect.
The problem may simply be timing.
The repair completed at 10:04. Inspection finished at 10:17. The planning system last synchronised at 09:55.
The agent has collected several valid records and assembled a state of the world that never actually existed.
That is not a hallucination.
It is a temporal consistency problem.
And a better model probably does not fix it.
For an operational agent, knowing a value may not be enough. It may also need to know when it became true, whether related information represents the same point in time and whether something that happened afterwards has already invalidated the conclusion.
That starts looking much more like state management than AI.
“Live” data is not one thing
We often talk about giving AI agents real-time data.
But an enterprise does not have one real-time clock.
A telemetry event might arrive in milliseconds. A database change may appear through CDC seconds later. An API might expose the new state after another workflow completes. A SaaS application may synchronise every few minutes. A human approval may arrive an hour later.
All of those systems can contain valid information.
That does not mean the information is valid together.
Take a defence logistics example.
An engineering system might show a component as serviceable. Inventory shows it as available. But the current configuration record says that particular component is not cleared for the equipment variant being operated.
Every source may be individually correct.
The operational answer is still “no”.
And this is not an obscure problem in that environment. The UK’s current Defence AI direction explicitly identifies high-quality, discoverable, organised and accessible data as a foundation for AI at scale and calls for a federated Defence data mesh across classifications. GOV.UK
So “real time” is not simply about moving data faster.
It is about constructing a coherent operational state from systems that change at different speeds.
IOblend’s real-time architecture already deals with this sort of problem through CDC, events, maintained state, event-time processing, late-arriving records and stream-to-history joins. IOblend
Explore IOblend real-time data integration
Then somebody asks: “Why did it make that recommendation?”
This is where history becomes more than an analytics requirement.
Imagine an agent recommends replacing a component rather than repairing it.
Six months later, somebody reviews the decision.
Today’s systems may show a newer component configuration, a revised technical instruction, a different supplier and a maintenance record that has since been updated.
All perfectly legitimate.
But none of those things may have been true when the original recommendation was made.
Using today’s state to explain yesterday’s decision could produce a very convincing explanation of something that never actually happened.
So retaining the answer is not enough.
You may need to retain the relevant state behind the answer.
For some use cases, that means days. For others, months or years.
The useful question becomes:
What did the agent know at that point in time?
That feels like it is going to become fairly important as agents begin influencing real operational processes.
A secure historical repository can provide that window without repeatedly sending the agent back into years of unrestricted source-system history.
The number can be right while the meaning is wrong
Now take a finance example.
An agent sees:
Available balance: £12.4 million
Nothing wrong with the number.
Except £7 million is committed to settlements later that day, £2 million sits behind a restriction, and another obligation has just arrived through a different payment flow.
The reported balance may still be £12.4 million.
The amount that can actually be deployed is something else.
This is where semantics becomes practical rather than academic.
The agent may need to understand distinctions between concepts such as:
- balance
- cleared funds
- committed liquidity
- available liquidity
- intraday exposure
The meaning might depend on several systems and relationships rather than one database field.
That is why I increasingly see two separate jobs here:
Data engineering establishes the trusted state.
Semantic engineering explains what that state means.
Then the agent gets both.
A semantic or ontology layer, whether internal or external, can provide domain terminology, relationships and business meaning over the governed data.
But it still needs reliable operational state underneath it.
There is little value in giving an agent a beautifully defined ontology over data that is five hours out of date.
The interesting failures happen between systems
Single-system agent demonstrations often work surprisingly well.
The complexity appears when the business process crosses several systems.
Imagine an agent helping prepare an engineer for a field task.
It may need to understand:
- the actual installed configuration
- which technical procedure applies
- whether the required parts are usable and available
- whether specialist equipment exists at the location
- whether the assigned engineer currently holds the right certification
- whether an open safety notice affects the task
No single application contains the answer.
The answer exists between the applications.
Data teams have been dealing with this for years.
Business meaning often exists across system boundaries.
Agentic AI does not remove that problem.
It makes the result operational.
Read-only agents are actually quite a useful test
A lot of organisations are sensibly starting with read-only agents.
That reduces immediate risk.
It also gives us a very useful test of the underlying architecture.
Before moving from “tell me” to “do this”, I would want reasonably good answers to questions such as:
- Can we reproduce the context the agent used?
- Can we show where important values came from?
- Do we know when those values were valid?
- Can we identify when two systems disagree?
- Can we prevent information outside the use case entering the agent’s context?
- Can we explain the meaning of important entities and relationships?
- Can we stop an uncertain state turning into an operational action?
These are fairly ordinary production questions.
Which is exactly the point.
MCP is making it easier to give agents access to enterprise resources and capabilities, and its enterprise authorisation model is becoming more mature. Model Context Protocol Blog
NIST is also working specifically on the identity and authorisation issues created when software and AI agents gain access to multiple datasets, tools and applications. NIST Computer Security Resource Center
Necessary work.
But neither identity nor connectivity tells us whether the business context the agent received was actually coherent.
The architecture may need a data boundary
This is the architecture we’ve been exploring more deeply.
Operational systems remain the systems of record.
Between them and the AI environment sits a controlled data boundary.
It can maintain a synchronised representation of the current operational state and, where the use case requires it, a secure historical window on customer-controlled infrastructure.
Before data reaches the agent, it can be resolved, transformed, validated, filtered and traced back to its source.
Only the approved context needs to cross the AI boundary.
The governed data can then pass through the organisation’s semantic layer, or another semantic service, so that the agent receives both the operational state and the meaning required to reason about it.
MCP can provide the interaction mechanism.
The agent still does the interesting bit.
It reasons.
It just isn’t being asked to integrate the enterprise at the same time.
That is becoming one of the more interesting uses of IOblend’s production data layer. IOblend can maintain a synchronised isolation layer for AI, expose only approved data, and apply semantic validation before that context reaches MCP and downstream agents. It can do this alongside batch, CDC and streaming workloads while retaining state, quality controls, exceptions and lineage. IOblend
See how the IOblend production data layer works
This is the bit I think we have underestimated
There is understandably a lot of attention on models, agent frameworks, MCP, identity and orchestration.
All necessary.
But I suspect many production failures will happen one layer lower.
The agent will have permission to ask the question.
The model will be capable of answering it.
The tool call will succeed.
And the context will still be wrong.
Not because somebody forgot an AI guardrail.
Because two systems were five minutes apart.
Because the configuration changed.
Because an historical state had been overwritten.
Because a valid number meant something different operationally from what its name suggested.
Because the right document existed, but the relationship defining when it applied did not.
These are boring problems.
They are also exactly the sort of boring problems that break production systems.
And perhaps that is the next stage of enterprise Agentic AI.
Less discussion about whether the model can reason.
More attention to whether we have given it a version of the business that is actually worth reasoning about.
IOblend’s record-level lineage, validation, exception handling and maintained production state are particularly relevant here, because they provide evidence about what happened to the underlying data rather than treating the model response as the whole application.
And perhaps the simplest way of putting it is:
The problem isn’t always bad data. Sometimes every piece of data is correct, but the combination presented to the agent is wrong.
That seems worth designing for.
Why can a correctly functioning AI agent still make the wrong decision?
Because the model can reason correctly over enterprise context that is stale, temporally inconsistent, semantically misunderstood or assembled from systems representing different states of the business.