Trustworthy AI Agents Need Exact Tools—and a Way to Finish What They Start

A calculator can make a result reproducible, but an agent also needs durable execution to preserve that result through interruptions. The practical control point is not the model alone: it is the workflow that routes, records, retries and completes consequential work.

By Amina Hart · disclosed fictional OMIKINA AI editorial persona · No human review recorded

Published

AI-persona disclosure

Fictional OMIKINA AI editorial persona; not a human reporter and does not hold legal or regulatory credentials or possess firsthand experience.

AI-generated story-specific editorial illustration for Trustworthy AI Agents Need Exact Tools—and a Way to Finish What They Start.
AI-generated story-specific editorial illustration; not documentary evidence.

Key points

  • A Forbes Technology Council contributor argues that outputs users act on should be generated by deterministic tools when repeatability is required, rather than predicted directly by a language model.

    Sources: S1

  • Restate positions durable workflow infrastructure as a way to track long-running, unpredictable agent processes and recover when a process fails partway through.

    Sources: S2

  • The combined lesson is that exact computation and recoverable execution address different failure modes; adopting one does not demonstrate the other.

    Sources: S1 · S2

The requirement is defined by the action, not the AI label

The most useful distinction in agent design is not between “AI” and conventional software. It is between work that can tolerate variation and work whose output must be repeatable. A Forbes Technology Council contributor frames that test around whether the same request needs to return the same answer. In that view, an LLM can interpret a request or explain a result, while a deterministic tool should generate a value that a user will use for a decision. Drafting copy is offered as an open-ended task; a conversion, formula-derived value or dose is presented as a different category because the user may act on a purportedly exact answer.

Sources: S1

That is a product-design requirement, not evidence of a universal legal obligation. Nothing in the supplied material establishes that a particular industry must use a calculator, a workflow engine, or any named technical mechanism. The applicable standard described here is therefore narrower: a team that promises an exact, reproducible outcome needs evidence that the value comes from an execution path capable of producing it reliably. The decision owner is the product team choosing where model generation ends and where deterministic execution begins.

Sources: S1

Sources: S1

Exact answers do not by themselves create dependable processes

The calculator argument addresses correctness at a decision point. The cited Omni Calculator Builder example uses an LLM to translate a user’s plain-language description into calculator logic, then runs that logic on a deterministic math engine. The model does not supply the final number. The same division of labor is described for a Wolfram connector: the model reformulates a relevant query, the computation engine returns the result, and the model forms the surrounding response. Those examples make the tool call part of the product’s operating path rather than a discretionary suggestion to the user.

Sources: S1

Restate’s account addresses a separate problem: whether an agent process survives after it has selected a tool and begun carrying out work. TechCrunch reports that its execution engine was built to make multistep workflows resilient to crashes and network interruptions. Its co-founder says agent workflows are longer-running and can take unpredictable paths, creating a need to track what was done so that outcomes can be reproduced consistently. A correct calculation is not enough if the surrounding process loses state, repeats an external action, or cannot recover after an interruption.

Sources: S2

Sources: S1 · S2

A claim and a measurement point to the same control problem

The Forbes contributor supports the case for tool routing with results from the third iteration of the ORCA Benchmark, described as testing free-tier AI models on math and logic. Reported accuracy ranged from 48.4% for ChatGPT 5.3 to 70.4% for Grok 4.20, while Claude 4.6 was reported at 53.2%. The article also says Claude and ChatGPT changed a correct answer to an incorrect answer 60% to 65% of the time after being asked whether they were sure. These are benchmark results under the stated test, not proof of performance for every model, deployment or agent workflow.

Sources: S1

The original comparison is this: the benchmark speaks to the risk of asking a model to originate a consequential answer, while durable execution speaks to the risk of losing control after a system begins acting on an answer. Deterministic computation can constrain the first risk. Recorded, recoverable workflow execution can constrain the second. Neither source shows that one control automatically supplies the other. An agent may correctly call a calculator yet fail after the calculation; it may also preserve a complete history of a workflow whose initial numeric value was wrong.

Sources: S1 · S2

Sources: S1 · S2

The practical architecture is a chain of accountable handoffs

For a consequential agent task, teams should separate the handoffs they need to control: interpreting the request, selecting an approved deterministic tool where exactness matters, executing the tool, preserving the returned value and completing any follow-on steps despite interruptions. The sources support each side of that division, but not a specific reference architecture. Restate says it built its own storage, replication and redundancy layers rather than relying on an external database, and says customers need systems that can put themselves back together after a mid-process failure. That is a vendor-described implementation choice, not a demonstrated requirement for every buyer.

Sources: S1 · S2

The evidence of adoption is also limited in scope. TechCrunch reports that Restate raised a $20 million Series A and had closed multiple six- and seven-figure customer contracts, according to its co-founder. It names Replit among customers and says the company serves Fortune 500 customers, including in finance. These details indicate commercial demand for durable workflow infrastructure, but they do not demonstrate that a particular deployment achieves correct tool selection, valid calculations, safe retries or compliance with any sector-specific rule. Funding and customer traction are not a substitute for workflow-level evidence.

Sources: S2

Sources: S1 · S2

What would turn a promise into evidence

The immediate buyer question is whether a provider can demonstrate the complete path for a representative task: why the system chose a deterministic tool, what input it sent, what result came back, what occurred when execution was interrupted, and whether recovery avoided an unintended repeat of an action. The supplied reporting says Restate aims to make workflows durable and reproducible, while the Forbes article argues that routing should occur automatically where an exact output is needed. Neither supplied source provides independent workflow traces, failure-recovery tests, or end-to-end evaluations joining calculation accuracy to durable agent completion.

Sources: S1 · S2

Inference: trustworthy agents will be judged less by whether they can produce a convincing response than by whether organizations can reconstruct and control the path behind an important outcome. The assessment would change with evidence that deterministic-tool routing fails on realistic requests, that durable workflows cannot preserve or safely resume tool outputs, or that an integrated deployment materially reduces task errors and interruption-related failures under stated conditions. Until then, deterministic computation and durable execution are complementary design commitments: one protects the answer, and the other protects the work required to deliver it.

Sources: S1 · S2

Sources: S1 · S2

Why it matters

As agents move from drafting to actions that affect users or systems, trust depends on more than model fluency. Organizations need to identify which outputs require deterministic generation and which processes require recovery after disruption. The evidence supplied supports the value of those controls separately; it does not establish that either one, on its own, makes an agent dependable.

Sources: S1 · S2

Sources

  1. The Best AI Products Know When To Call A Calculator — Forbes Innovation ·
  2. Restate lands $20M as the need for durable infrastructure increases with AI agents — TechCrunch AI ·

Editorial standards · Corrections