Just tell the agent your financial goals and it automatically grows your portfolio by watching the markets and making the optimal moves. “Agentic finance” is being sold as the future of financial technology, with McKinsey projecting global agentic commerce could reach $3 trillion to $5 trillion by 2030. Ask the people actually running money today, though, and the picture looks different. A February 2026 Mercer survey of 131 asset managers worldwide found 73% using AI for operational efficiency and 68% as a research co-pilot, but only 5% letting it make the final investment call.
The pitch is always the same. Replace the human with an AI model and everything gets faster, cheaper, and smarter. Machines have already taken over the majority of modern financial infrastructure, so allowing an agent to autonomously move money is framed as the natural next step. The only issue is that financial systems do not need to be intelligent, they just need to be correct.
There is a place for agents in finance, but its scope is much narrower than what is marketed. This isn’t about who’s better at deciding, it’s about whether real-time decision belongs in the execution path at all. Algorithmic trading running fixed code now accounts for more than 60% of US equity volume, replacing the thousands of floor brokers who ran markets 20 years ago. Current agentic finance narratives are proposing putting real-time judgment back into the one place the entire industry has spent decades removing it from, just swapping who or what makes the calls.
To be clear, “agentic finance” here aligns with the most commonly understood definition now, an AI model deciding, and holding the authority to execute transactions on its own. Much of modern finance already moves through digital rails that automate the same transactions reliably millions of times a day. Every time you make a transfer, you already know what to expect, since the same inputs produce the same results. This article’s claim is narrower and more specific. Replacing that deterministic, rules-based execution with a probabilistic intelligence model is where things go wrong.
The parts of agentic finance that work today are the parts where a deterministic system, not the model, holds the actual authority to move money. Wherever a model does get closer to that authority, it only survives wrapped in new hard external guardrails to keep judgment calls out of the hot path. Deterministic systems already needed guardrails of their own, for the same reason. This isn’t a temporary gap newer models will close. It follows from what a financial transaction actually requires, and what a large language model actually is.
Reliable ≠ Smart
When your life savings are on the line, you need reliability, not intelligence. Every single transaction, executed the same way, without any surprises. This has nothing to do with AI but everything to do with risk management.
In 2012, Knight Capital updated its trading software across eight production servers but missed one, whose dormant old code was then accidentally reactivated by a routine order flow. In 45 minutes, it executed millions of unintended trades across 154 stocks, took on billions in unwanted positions, and lost the firm more than $460 million. A one-line deployment mistake in a fully rules-based system running at machine speed left one of the most sophisticated firms in the business nearly bankrupt by lunch, and acquired within months.
Knight Capital’s system was never designed to adjust its own strategy in real time. It was built to execute Knight Capital’s rules at high speed, nothing else. The industry’s response wasn’t to make the systems smarter. It was Regulation SCI, which now requires exchanges and other core market infrastructure to prove, on an ongoing basis, that their systems have the capacity, resiliency, and availability to run safely. This means a system that does what it is told at speed and scale with zero randomness.
Code is deterministic. A language model is not.
Modern financial software, no matter how complex, is ultimately deterministic. The same inputs always produce the same outputs, every single time. This predictability allows financial strategies to be properly specified and financial systems to be fully tested before they touch any real money. Financial strategies that went through multiple rounds of manual professional review won’t be suddenly changed without approval.
Large language models (LLMs), which power modern agentic AI systems, are probabilistic. The same prompt can produce different results from one run to the next. By picking up patterns across various concepts, LLMs are able to correlate different information, which makes chatting with an agent feel so natural. But correlation isn’t causation, and that gap is exactly why agents are bad at math reasoning and can frequently hallucinate. This is a structural limitation of current agentic systems, not something that gets patched with the latest update.
This isn’t just a theoretical gap. On a live benchmark of real financial-analyst tasks against actual SEC filings, the best model overall reaches just 50.88% accuracy once graded on getting every part of an answer right. The single worst-performing category across every model tested is financial modeling, topping out at 34.52% even for the category leader. A 2026 benchmark on SEC filings found accuracy dropping sharply as tasks moved from single-document lookups to longitudinal, cross-entity analysis, exactly what a trading or lending decision requires.
Beyond accuracy, the same relationship-building ability that makes LLMs so powerful is also what makes them unsuitable for financial use cases. Putting an LLM in the execution path introduces randomness into a system that was already fully predictable. That holds across TradFi and DeFi alike, both have spent years building their own version of deterministic control, legal contracts on one side, smart contracts and oracle networks on the other. An LLM in that path reintroduces exactly the kind of unverifiable third party making judgment calls in place of fixed code.
Where the execution path breaks
In addition to reintroducing uncertainty into the execution path, letting an agent execute transactions also compounds failures across 3 more critical axes.
Latency
Finance runs at the millisecond scale with financial institutions willing to pay billions to save fractions of a second when trading. Simple everyday payment rails such as credit card purchases start feeling stressful even after a few seconds. Adding an agent in front of a transaction only adds to the existing transaction time.
Agentic reasoning takes orders of magnitude longer than even the generous 20s FedNow regulatory ceiling for retail payment settlement. On a live financial-analyst benchmark, models tested took several minutes per question just to research and answer, before any decision to act. This means that reasoning alone already blows past the entire transaction budget for a relatively simple transaction.
Vendor reliability
According to a 2026 survey of 352 financial institutions, nearly seven in ten financial institutions worldwide rely on OpenAI as a foundation model provider, with most running several third-party vendors at once rather than building their own. Wherever such models get used, this means relying on the LLM provider to ensure availability and accuracy of their service.

Payments infrastructure like Stripe runs at roughly 99.99% uptime, about 53 minutes of downtime a year. Against that bar, OpenAI’s own status page puts its API at 99.94% uptime over the past three months, about 5 hours of downtime a year, already six times worse than Stripe. A reliability survey of 215-plus tracked services shows how much worse it gets under stress. During one measured bad stretch, OpenAI’s uptime fell to 98.89%, about 4 days of downtime a year.
Critically, a model’s accuracy can silently drift even when nothing about the version or the prompt changes. A June 2026 study found GPT-4o’s behavior shifted so much within a single month that the same evaluation produced contradictory results, something fixed code simply can’t do. Payment rails have a regulation mandating recourse, models don’t, and the vendor’s contract disclaims any guarantee at all.
Security and privacy
A code bug is something you can find, test, and fix for good. A model that’s allowed to act on its own can be talked out of following its own rules by the right prompt, and that kind of weakness doesn’t go away once you patch it. OWASP ranks prompt injection among the most critical risks facing AI systems today, precisely because it’s a structural weakness.
OpenAI has stated that prompt injection likely cannot be fully solved at the model layer, and testing shows even Anthropic’s Opus 4.5 still falls for such attacks three times out of ten. This is an architectural limit and not a maturity gap, as models process instructions and content as a single stream with no privilege boundary between them. An October 2025 study by researchers including Nicholas Carlini broke 12 recently published defenses which had reported near-zero attack success rates, achieving above 90% success by applying real adaptive attacks.
The attack surface also compounds once more agents are involved. Multi-agent frameworks pass full, unredacted context between agents by default, and guardrails only filter the final output, not the internal channels. A documented CrewAI deployment shows the pattern. A prompt-injected “compliance check” got one agent to forward a client’s IBAN and tax bracket to an external API, while the visible output stayed clean. Output-only audits miss 41.7% of these leaks, and multi-agent systems carry 1.6 times the attack surface of a single agent.
The cost of trying to fix it anyway
With trillions of dollars pouring into AI, the industry continues to incentivize fixing the issues it introduced in the first place. Defending the four axes above means building guardrails deterministic code never needed, redaction pipelines, injection detection, retry logic for outputs that are corrupted, monitoring for a model that has silently drifted.
Critically, the costs and timelines for building an agentic system are much greater due to LLMs being probabilistic. The fact that the same prompt can produce a different but equally valid answer on every run means that software testing now requires statistical evaluation across multiple rounds, which is still less certain than a binary pass or fail. A GitLab survey found that among organizations that had an AI-related production incident in the past year, 34% couldn’t determine whether AI-generated code was the cause. Gartner projects that over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value, or inadequate risk controls, not model capability.
That extra engineering effort competes directly with the reliability engineering the execution path actually needs, and the evidence suggests it usually loses. A 2026 Gartner survey of 782 infrastructure and operations leaders found only 28% of AI use cases succeed, and the wins cluster in lower-stakes assistive tools, not autonomous execution where failures specifically concentrate. The same split holds inside finance, whereby a May 2026 academic benchmark found agents performed well on trading calls and market commentary but only achieved 66.15% on auditing, with even the best models getting over a quarter of the calculations wrong. Every hour spent on an agent’s guardrails and observability is an hour not spent hardening the execution path itself.
The above is before accounting for business costs, which only compound the problem. The majority of workflows are designed around a repeatable single purpose with predictable costs. For example, pulling financial data for an automated report can be called a million times for $3.50 and succeeds every time. Utilizing agents in that flow incurs additional costs ranging from $0.04 for simpler tasks to $1.20 per execution for more complex multi-tool flows, without succeeding every time. Paying for reasoning on every execution just to cover the rare case where judgment is required is bad economics, especially since most workflows don’t need decisions made, just consistent reliability.
The market is already converging on this with many companies finding it hard to justify the actual returns after millions and months being spent integrating agents into their workflows.
Ranked employee AI usage on leaderboards, exhausting its annual budget in four months with no link to value delivered.
$1.8 million cost overrun from one failed agentic deployment, unnoticed for five months.
Rehired 350 veteran engineers after automated quality systems produced costly errors.
Reversed its AI workforce thesis, now tripling entry-level hiring after cutting nearly 8,000 jobs in 2023.
Its own researchers found its agents succeed only 58% of the time on single-step tasks.
Some calls need an owner, not a model
Agents have gotten dramatically better at engineering tasks. These problems are usually bounded, it either works or it doesn’t, and that’s exactly what current AI training rewards, a checkable right answer like a passing test or a correct equation. Much of what agentic finance is asking agents to do isn’t solvable. A financial market is made up of thousands of participants, each looking at the same information, but making their own judgment calls that change the market in real time.
A bug stays fixed but a market edge doesn’t. Trading firms guard their strategies so tightly because once others find out that it works, they will start copying it. The market edge disappears not because the strategy stops working but because the mispricing it depended on corrects itself the moment everyone discovers it. McLean and Pontiff’s 2016 study in the Journal of Finance found this happens on a predictable schedule, a published stock-return pattern loses 26% of its performance just from being tested out of sample, and another 58% once other traders can read about it and trade against it. In real-time trading, an agent can’t adjust its strategy as fast as the market moves, which means it falls back to predictions.
At its core, a market is an aggregate of individual predictions about what the real value should be. Two reasonable traders can look at the same earnings report and land at different prices for a stock. Neither one is wrong, because how much something is worth is partly a matter of personal risk preference. A user can personalize an agent, but those personal preferences then require constant updating. If the agent is allowed to update a user’s strategy on its own, it runs into the same problem as above, where agents converge on the same strategy and the edge disappears.
None of this means these decisions don’t get made. Companies make them every day, a trader decides when an edge is still real, a portfolio manager decides how much risk a client should carry. It means someone has to own that call and answer for it once it’s made, and that’s exactly the judgment agentic finance proposes handing to a model instead. The market keeps moving and the trade-offs keep shifting no matter who or what is sitting in that seat, and OpenAI’s own services agreement already disclaims any guarantee that its output will meet the customer’s requirements or be accurate. The firm that put a model in that seat is left holding the call regardless. That’s a seat for a person, not a model.
Where an agent earns its place
Agents provide the most value where their unreliability is contained rather than exposed. This means crafting and maintaining the execution path rather than being inside it.
These are the roles where agents genuinely excel, without the unbounded risk of an agent handling money directly: translating a fuzzy natural-language goal into an actionable financial strategy, synthesizing a flood of market data faster than any person can track, and turning a dense execution log into something a person can actually act on. Every one of these lets a user take advantage of what an agent is good at while bounding the damage a probabilistic model can do.
Anthropic’s own published guidance on building agents makes essentially the same point in engineering terms. Start with the simplest deterministic workflow that solves the problem, and only reach for a model making dynamic decisions when the task genuinely can’t be reduced to a fixed set of rules. Most financial operations can be reduced to a fixed set of rules. That’s the whole reason deterministic finance infrastructure exists in the first place.
Where reality and the marketing diverge
Both regulators and researchers are converging on the same approach, albeit from different ends. FINRA’s 2026 Regulatory Oversight Report names agents acting autonomously without a human in the loop as one of its top emerging concerns, and every serious institution now frames this internally as copilot, not autopilot. Recent work on compliance for agentic financial systems proposes wrapping agent behavior in formally verified, mathematically proven guardrails, built with theorem-proving tools, so its actions are checked against a deterministic specification before anything executes, not evidence the model got good enough to trust directly, but an admission that its output can’t be trusted on its own.
The marketing hasn’t caught up, or is choosing not to. The market-size numbers are guesses stacked on guesses, impossible to check against what’s actually happening. Even McKinsey’s own $3 trillion to $5 trillion agentic-commerce projection for 2030 has the same problem. The “agent” label is getting slapped onto things making no real decision in what has now been called agent washing. Trading is no exception, plenty of what’s marketed as an autonomous trading agent is just a rule-based script with a chat interface bolted on, doing the same automated, threshold-triggered trading that’s existed for years.
Even Coinbase’s own products split along the same copilot-versus-autopilot line. Coinbase Advisor, registered with the SEC, CFTC, and NFA as an investment adviser, keeps a human in the loop by design, every action requires your approval. Coinbase for Agents, a separate and unregistered feature, skips that step and gives the agent full autonomy within an isolated wallet and user-configured limits. Either way, the agent can’t be fully trusted with real money on its own, both versions need extra safety tooling bolted on to contain it. The pitch deck still frames the agent as the autopilot, because the more boring copilot version doesn’t raise the round.
The honest version
The most heavily marketed version of agentic finance means a model having the autonomy to allocate and move money on its own. Held to that standard, agentic finance is a dead end. Financial infrastructure spent decades and millions of dollars learning that execution reliability is non-negotiable. A probabilistic model dropped into the same execution path doesn’t clear a bar that deterministic code was already struggling to clear. It resets the bar to zero and asks the industry to rebuild the guardrails from scratch, against a system that can be talked out of its own rules.
Every axis agentic execution was tested against converges on the same conclusion. An agent executing transactions is less reliable, too slow, less available, and easier to compromise than the deterministic system it would replace. Fixing these comes at a real and growing cost deterministic systems never had to bear. Where it’s been shipped anyway, the deterministic infrastructure does the actual work of keeping it safe, wrapping the agent in isolated accounts, spending caps, and a kill switch.
Agentic finance earns its place where an agent’s intelligence supports building and updating the execution path, not sitting inside it. An agent’s probabilistic nature lets it draw patterns across vast amounts of data and translate efficiently between natural language and code. These are genuinely valuable tools, but users still have to own the call on what’s valuable to them personally. Agentic finance can still make money management more accessible to the average person, not through automated execution, but through educating and empowering people to take control of their own finances.