The industry has a word for this problem, and it is pointing the word at the wrong thing. Determinism, in AI, means the same question gets the same answer every time, which is what a regulator expects and what a probabilistic model does not naturally give. Most of the effort in financial infrastructure right now goes into making the model deterministic. Agents are arriving before the harder version of the problem has been solved. EY found that 31 percent of the banks it surveyed in 2025 had started implementing agentic AI. In the same year, a Harris Poll of more than 500 U.S. financial decision-makers found that 98 percent still perform some payment operations manually and 49 percent use five or more systems to manage payments. Institutions are preparing software to make decisions and take action across operations that still depend on people to interpret what was supposed to happen. That is the constraint, and it is not the one the industry is arguing about.
Getting the same answer from an AI model every time is the easier problem
There is now a familiar architecture for putting probabilistic models into consequential workflows. Let the model handle ambiguity. Turn stable, repeatable decisions into explicit logic. Define what the agent is allowed to do, log its actions, check the results, and escalate anything outside its authority. The principle is sound: do not ask a probabilistic model to rediscover the same answer thousands of times when the institution already knows what the answer should be.
Governance is moving the same way. In February 2026 the Treasury released a Financial Services AI Risk Management Framework, developed with more than a hundred institutions and aligned to the NIST AI framework, with 230 control objectives that turn accountability, transparency and oversight into things an examiner can test.
All of it starts downstream of a more basic question: what was supposed to happen? A rule that fires the same way every time is no use if the expectation behind it is wrong, incomplete, out of date or sitting in somebody's head. You can solve the consistency of the model without solving the consistency of the operation. A deterministic model running on an undocumented operation is deterministic about nothing that matters.
The documents describe the operation. They do not run it.
Take every contract, policy, procedure and network mandate an operation runs on, hand the stack to a capable new hire and ask them to run settlement for a week. They still cannot. A processor agreement says T+2. An amendment moves one product to T+1, and a holiday adds a third timing rule. The fee schedule says one rate, but an annex changed it in March. The written procedure assigns an exception to one team, and the operation has been routing it somewhere else for six months because that is where it gets resolved. None of those facts necessarily lives in the same system, and having the documents does not tell you which expectation governs the transaction in front of you.
The operation runs on a hierarchy of contracts, amendments, policies, procedures, network mandates, system states, approved exceptions and accumulated operating knowledge, and people have always been the layer that holds it together. They remember which amendment superseded which schedule, which processor code means what, which partner the written procedure quietly does not apply to, and who to call when two systems disagree. That is why an operation can look highly consistent and still be very hard to automate. People are absorbing the ambiguity. An agent cannot inherit that knowledge by connecting to the applications. This is the wall every enterprise AI program hits when a prototype meets production, and it has little to do with the model. An agent cannot reason about a business it does not understand: what the data means, how the teams actually work, and how this company defines a word the rest of the world defines differently. In one system revenue means gross, in the next it means net, and nobody wrote down which one the settlement report uses. In financial operations that missing context has a specific shape. It is the expectation the transaction should have met, the definition each system is using, and the person who owns the difference when the two do not agree. Context has to come before automation, and the part of context that can be written down is the operating rulebook.
A two percent exception rate can hide the real problem
Operations leaders already have a number for how well this works: the exception rate. An exception is a case the normal path could not settle on its own. A settlement that landed a day late, a fee that did not match the schedule, two records of the same payment that disagree. It drops out of the automated flow and goes to a person to resolve. A low exception rate is how a well run operation shows it is in control, which makes it the natural answer to the argument above.
So the most common objection comes from well run operations, and it is a fair one. Our exception rate is two percent. The team clears it. Where is the problem?
The two percent is not the interesting number. The question is what lets the other ninety-eight percent run cleanly. If it runs because the operation contains an explicit, testable expectation, that work may already be ready for more automation. If it runs because experienced operators resolve small ambiguities without thinking about them, the workflow is less deterministic, meaning less certain to be handled the same way every time, than its exception rate suggests.
A person knows that two slightly inconsistent records refer to the same payment, that this counterparty's T+2 behaves differently across a holiday, that the amendment signed in March overrides the fee schedule still sitting in the shared drive. An agent either receives that context or has to infer it, and an inference, however good, carries no authority.
That is the real readiness question: how much of the normal path can be explained without relying on unwritten human judgment?
Test one workflow: how many of its rules could a new hire follow without asking anyone?
Pick one workflow you would actually hand to an agent: partner settlement, merchant onboarding, interchange assurance, fee validation or returns. Walk it decision by decision and, for each one, ask what should be true, under what conditions, for which counterparty, product, rail or jurisdiction, from what date, on the authority of which source, superseded by what, proven by what evidence, and handled how when the evidence is incomplete.
A rule that passes looks like this: settlement from this processor for this product lands T+2 on business days and T+3 when the settlement date crosses a recognized holiday, under section 4.2 of the January agreement as amended on March 14. Every batch is compared against that schedule, and anything outside it routes to settlement operations with the evidence attached.
A rule that fails looks like this: settlement normally arrives in a couple of days, and if it does not, Maria knows whether it is actually late.
Count how many consequential decisions pass. That is a far more useful measure of agent readiness than the exception rate, because it tells you how much of the workflow can be delegated without asking a model to invent the institution's policy as it goes.
An operating rule needs a source, a scope, a date, and an owner
Writing the condition down is only the beginning. A financial operating rule needs provenance: it traces to a contract, amendment, policy, procedure, network mandate, regulation or an explicitly approved operating decision. It needs scope, because the same expectation may not apply to every processor, merchant, jurisdiction, rail or transaction type. It needs time, because agreements change and the system has to know which version governed the event when it occurred, not which rule is current today. It needs precedence, so that when two authoritative sources disagree, something establishes which one controls.
And it needs ownership. When a rule is inferred from how the operation behaves rather than found in a document, observation alone should not turn it into institutional truth. The system can propose the expectation and show the evidence. A person accountable for the operation has to approve it. Observation proposes. The customer ratifies.
Not every judgment should become a rule
The goal is not to turn every judgment a model makes into permanent logic. Some cases are contextual, some evidence is incomplete, some situations are new, and some policies deliberately leave room for judgment. Those belong with a person, a model, or both.
The rulebook should take the stable part of the operation off the table: the obligations, thresholds, timing windows, allowed states, calculations, eligibility criteria and escalation conditions the institution already knows. Everything else should show up as uncertainty rather than being forced into false certainty. That gives an agent three possible answers: the evidence confirms the expectation, the evidence does not confirm it, or there is not enough evidence or authority to decide. The third answer is not a failure. In financial operations, knowing when not to act is part of the control.
The context an agent needs before launch is deterministic truth about the operation
Put the pieces together and the requirement is plain. Before an agent can be launched into a financial operation, it needs a source of truth that answers the same way every time: what this transaction was supposed to do, under which agreement, as of which date, and who decides when it did not. That is deterministic truth in the only sense that matters to an operation, and no model can generate it. It has to be assembled from the documents, the definitions each system uses and the decisions people have already made, kept current as those change, and held where every agent can read it.
That is the context layer. It sits between the model and the systems. The model supplies the reasoning. The systems supply what happened. The context layer supplies what should have happened, with the authority behind it. Without it, an agent connected to every system is still guessing at the one thing it cannot see. With it, the agent's job becomes what it is good at: comparing, reasoning about the difference, and acting within the authority it was given.
Regulators are starting to ask what the agent was supposed to do, not only what it did
It is worth being exact about where regulation stands. In April 2026 the OCC, the Federal Reserve and the FDIC issued revised model risk guidance covering development, validation, monitoring, governance and controls, and said plainly that generative and agentic AI sit outside it because they are still novel and fast moving. The agencies also said further work on banks' use of AI is coming. So there is no rulebook test an examiner will demand next quarter.
The direction is visible all the same. The Treasury framework asks for lifecycle governance, accountability, monitoring, evidence and effective controls. As agents take on more consequential work, institutions will need to establish what authority an agent had, what information it used, what it did and when a person was required to step in. Underneath every one of those questions sits another: what was the correct outcome supposed to be? An execution log tells you what the agent did. A rule with provenance tells you why that action was right. You need both.
Where the command center fits
Cordant is the command center for modern financial infrastructure. It sits above the processors, banks, ledgers, compliance systems and counterparties an institution already runs, replaces none of them, and connects what happened with what was supposed to happen. The expectation comes from the material governing the operation: agreements, amendments, policies, procedures and approved rules. The actual comes from the systems and counterparties executing the work. Cordant compares the two continuously and attaches the evidence.
When the operation reveals a consistent rule that is not in the documented rulebook, Cordant proposes it with the evidence behind it. It does not quietly promote observed behavior into institutional truth. The responsible operator approves it first. Over time that builds the context layer an agent needs: a versioned operating map of the rules, the relationships and the actual behavior that determine how the financial operation works. A person can use it, a queue can use it, an agent can use it, and all three work from the same expectation.
What lasts is the institution's own definition of what should happen
The institutions that put agents into real financial operations will not get there because they chose a better model. They will get there because they built the context layer first. Models will keep improving and they will keep changing. What lasts is knowing what the institution expects to happen, where that expectation came from, which version applies, whether reality matched it, and who has authority when it did not. The model can reason. The agent can act. The institution still has to define what right means, and that definition, written down and checked against what actually happened, is the only determinism a regulator will ever ask about.
Money already moves in real time. The decisions should too.




