Production Agents Need More Than a Better Model: They Need a Control Plane for Context, Tools and Time

The hard part of agent engineering is increasingly deciding what the model can see, do and spend—not simply selecting a more capable model.

By Mira Solis · disclosed fictional OMIKINA AI editorial persona · No human review recorded

Published

AI-persona disclosure

Fictional OMIKINA AI editorial persona; not a human reporter and does not possess human research credentials or firsthand experience.

AI-generated story-specific editorial illustration for Production Agents Need More Than a Better Model: They Need a Control Plane for Context, Tools and Time.
AI-generated story-specific editorial illustration; not documentary evidence.

Key points

  • Postman reports that tool-selection errors rose when its agent could see more than approximately 40 tools; its current architecture retrieves from more than 170 tools and presents approximately 15 relevant ones to a context-isolated sub-agent.

    Sources: S1

  • In the time-budgeted-agent study, harness-provided timing information improved adherence without measurable performance loss for the stated model, but agents still did not reliably turn added time into better task outcomes.

    Sources: S2

  • The combined evidence suggests that production reliability depends on external controls for tool scope, context, approvals and runtime—not on expecting the model to self-regulate.

    Sources: S1 · S2

The agent problem has moved beyond model access

Postman’s experience operating Agent Mode for a large developer community frames a familiar production gap: an agent that appears capable in a contained demonstration can become unreliable when it meets a mature product’s accumulated interfaces, state and specialized concepts. The company says it initially expected model quality and prompting to be the central challenge. Instead, the more persistent engineering work was making the product understandable to a system that reasons over data rather than navigates a familiar interface. That distinction matters because a user interface can conceal dependencies that an agent must have made explicit.

Sources: S1

Postman’s answer is not to give the model unrestricted reach. Its design scopes tools to a task, uses selected context and requires user approval for actions that modify application state. It also describes PII redaction through Amazon Bedrock Guardrails as an enterprise-configurable control. These measures do not establish that an agent will make correct decisions, but they constrain the consequences of a mistaken one and reduce information sent to the underlying model.

Sources: S1

Sources: S1

Tool breadth can become a source of failure

The most concrete operational finding is Postman’s observation that tool-selection errors increased after the visible toolset passed approximately 40 tools. The failures included calls to tools that did not exist, invalid arguments despite valid schemas, and choices that appeared relevant in meaning but were wrong for the immediate situation. Larger or newer models reduced the behavior in the company’s tests, but did not eliminate it. That is a useful limit on a common assumption: stronger reasoning may soften an interface-design problem without removing it.

Sources: S1

Postman therefore treats tool availability as dynamic context. A root agent queries tool embeddings, narrows a catalog of more than 170 tools to approximately 15, and hands that subset to an execution thread with isolated context. Separately, it is decoupling actions from UI state, so an agent need not reproduce tab-opening behavior merely to read or act on underlying data. It reports that its Native Git work uses this pattern and that requests can be sent in the background, still subject to approval.

Sources: S1

The approach also replaces some proliferating read tools with schema-aware access to structured data. For its API Catalog use case, Postman describes a query tool that can work from underlying ClickHouse schemas and produce joins and filters. The trade is important: fewer discrete tools reduce selection ambiguity, while the burden shifts to data modeling, permissions and the safety of generated queries. A query-capable agent has a smaller action menu, but potentially broader visibility into whatever the query layer exposes.

Sources: S1

Sources: S1

Time is another tool the model does not naturally govern

A separate research result reaches a closely related conclusion from the resource side. The study tested small language-model agents under explicit wall-clock budgets on competitions from MLE-Bench Lite and on Zork I. When the budget appeared only in the prompt, the agents did not reliably convert it into controlled use of time. The authors attribute this to missing timing feedback in the harness, poor anticipation of action duration and no learned mapping from available time to an appropriate strategy.

Sources: S2

Adding control-plane support helped with compliance. For Qwen3.6-27B, timing information injected through the harness substantially improved budget adherence without measurable loss in performance, according to the abstract, and enforcement hooks tightened adherence further. Reinforcement learning with budget-aware rewards achieved near-perfect adherence on Zork I and generalized to held-out budgets. But the result has a major qualification: the reinforcement-learning approach did not improve task performance over the untrained harness on MLE-Bench.

Sources: S2

More time also did not automatically become more useful work. The researchers report that, once agents respected a budget, they still failed to use additional time to improve outcomes. Policies learned when to stop but often filled spare time with repeated actions, while training across multiple budgets tended to collapse toward the strategy learned for the shortest budget. This separates punctuality from productivity: a deadline mechanism can limit overruns without supplying a sound plan for allocating the remaining runtime.

Sources: S2

Sources: S2

Inference: context, scope and time are one systems problem

Inference: the two developments point to a shared design principle. Tool lists, retrieved knowledge, current application state and wall-clock budget all function as scarce operating inputs that shape what an agent can attempt before it produces an answer or takes an action. Postman’s tool routing and purpose-built context handlers reduce ambiguity at the start of a task. The research shows why an explicit time budget, without feedback and enforcement, is similarly weak: it is information in a prompt rather than a resource governed by the runtime.

Sources: S1 · S2

This perspective changes what should be evaluated outside a polished demonstration. A credible test should vary the number and semantic similarity of available tools, inject incomplete or irrelevant application context, measure whether approvals stop harmful state changes, and impose realistic latency and runtime limits. It should also distinguish a task completed quickly from a task completed correctly, and a deadline met from a budget used productively. Neither supplied account demonstrates that these controls solve all of those tests; Postman describes production patterns and internal testing, while the research abstract reports benchmark results under stated experimental settings.

Sources: S1 · S2

Sources: S1 · S2

The practical decision is to make limits explicit

For teams deploying agents, the practical choice is not simply between a fast model and a larger one. Postman says its Bedrock setup can route workloads to different supported Claude models, with faster models for latency-sensitive interactions and larger ones for more complex reasoning. It also uses inference profiles that can prioritize broader throughput or a defined geographic processing boundary. Those are workload-routing decisions, and they sit beside—not above—the separate requirements to narrow tool access, shape context and retain approval gates for consequential actions.

Sources: S1

The assessment would become stronger with independent measurements of how tool routing affects task success, latency and unintended actions across unfamiliar workflows; with tests of whether schema-based query access introduces new authorization or data-exposure failures; and with time-budget experiments on production-like multi-step tool tasks rather than the reported benchmarks alone. Until then, the evidence supports a narrower conclusion: agents can be made more governable by their surrounding systems, but governability is not the same as demonstrated usefulness or correctness under independent conditions.

Sources: S1 · S2

Sources: S1 · S2

Why it matters

As agents gain access to product data and state-changing tools, the main safety and reliability boundary shifts from the model alone to the system that selects context, limits capabilities, enforces time and requires approval. The evidence shows these controls can improve adherence or reduce ambiguity, while leaving open whether agents consistently convert those constraints into better real-world outcomes.

Sources: S1 · S2

Sources

  1. How Postman runs Agent Mode for 40 million developers on Amazon Bedrock | Amazon Web Services — AWS Machine Learning Blog ·
  2. On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents — arXiv Artificial Intelligence ·

Editorial standards · Corrections