Skip to content

PATTERN Cited by 2 sources

Deterministic tool vs LLM judgment

Deterministic tool vs LLM judgment is the discipline of deciding, for each unit of work in an agent system, whether it needs an LLM or a deterministic function — and implementing it accordingly. The practical test (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore):

"If I fix the input, will the output always be the same?"

  • Yes → implement as a @tool. The LLM decides when to call it and with what arguments; the function does the work and contains no inference.
  • No → the agent's LLM reasoning IS the logic. The variability is intentional (interpreting ambiguous instructions, contextual judgment, natural-language rationale).

The one-line principle: deterministic computation belongs in tools; judgment, interpretation, and context-dependent recommendation belong in agent reasoning.

Why it matters

  • Bounded cost. Wrapping a deterministic formula in an LLM adds cost, latency, and non-determinism with no benefit. Keeping inference to the handful of genuine judgment calls keeps per-run LLM cost flat as the workload scales.
  • Reasoning quality. Each agent's context carries only what it needs to reason about, not the raw byproducts of every tool call.
  • Testability & auditability. Deterministic tools are unit-testable; their outputs are reproducible and go straight into a structured JSON contract.
  • Prevents scope creep. When a new requirement arrives ("add a second validation check"), the answer is clear: add a @tool, not a new agent.
  • Failure containment. Asking a deterministic function to interpret "there's a promotion next week" will fail; asking an LLM to compute a replenishment formula invites hallucinated arithmetic. Matching each task to the right executor removes both failure modes.

Worked application (13 components → 6 agent, 7 tool)

Work Deterministic? Implementation
Parse natural-language request No agent reasoning
Load file from S3 Yes @tool
Select covariates from data quality + context No agent reasoning
Invoke Chronos2 endpoint Yes @tool
Interpret forecast anomalies in context No agent reasoning
Calculate order quantity from formula Yes @tool
Check order vs warehouse/budget constraints Yes @tool
Generate order rationale No agent reasoning
Generate chart / write JSON to S3 Yes @tool
Retry-with-adjusted-constraints vs escalate No agent reasoning
Summarize in business language No agent reasoning
Score forecast accuracy vs actuals Yes code-based evaluator

The second boundary: in-process vs Gateway

The determinism test decides tool vs agent. A second, orthogonal test decides where a tool lives — it turns on trust, durability, and cost-of-mistake, not determinism:

"If the agent hallucinates and calls this tool wrongly, does the mistake propagate to external systems or stop at the agent's memory?"

  • Stops at the agent → in-process @tool. IAM role + type system already bound it; adding a gateway adds latency/cost with no safety gain.
  • Propagates externally → Gateway with authorization (JWT identity + Cedar policy) and an audit trail of "who asked for this write, and what was persisted."

In the canonical instance, exactly one of eight tools (save_decision, the authoritative order write) needs the Gateway + Cedar policies; generate_forecast_chart also writes S3 but stays in-process because a chart is a visualization, not a decision of record. "That proportion is the norm, not the exception." Placing a high-value deny policy on the write boundary (not on an upstream calculate step) is what makes it meaningful — the agent could re-run a calculation until it passed, but the durable effect happens only at the write.

Extends to evaluator design

The same logic applies one layer up, to which evaluator type to use (see two-layer evaluation): forecast accuracy is arithmetic with a correct numeric answer → code-based evaluator; agent behavior (did it flag the violation, was the rationale coherent) is subjective → LLM-as-a-judge. "Use the evaluator type that matches the question, not the one that feels more sophisticated."

Relationship to other patterns

Seen in

  • sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore — the organizing thesis of the four-agent inventory pipeline; Bedrock inference happens in exactly the four agents, everything else runs as deterministic @tools, and only save_decision crosses the Gateway boundary.
  • sources/2026-09-28-redpanda-your-ai-kill-switch-is-in-the-wrong-place — applies the determinism discipline to authorization: a policy decision at the agent boundary must be a deterministic typed evaluation, not the agent's (or a guard-model's) judgment. "The same request in the same context must get the same decision every time, with nothing left to chance or the agent's own understanding of its boundaries." This is the security-boundary instance of "judgment belongs to the LLM, deterministic evaluation belongs outside it" — the allow/block/hold-for-human verdict is exactly the kind of reproducible, auditable decision that must not live inside the non-deterministic agent (Redpanda's OBPE).
Last updated · 766 distilled / 2,225 read