PATTERN Cited by 2 sources
Deterministic tool vs LLM judgment¶
Deterministic tool vs LLM judgment is the discipline of deciding, for each unit of work in an agent system, whether it needs an LLM or a deterministic function — and implementing it accordingly. The practical test (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore):
"If I fix the input, will the output always be the same?"
- Yes → implement as a
@tool. The LLM decides when to call it and with what arguments; the function does the work and contains no inference. - No → the agent's LLM reasoning IS the logic. The variability is intentional (interpreting ambiguous instructions, contextual judgment, natural-language rationale).
The one-line principle: deterministic computation belongs in tools; judgment, interpretation, and context-dependent recommendation belong in agent reasoning.
Why it matters¶
- Bounded cost. Wrapping a deterministic formula in an LLM adds cost, latency, and non-determinism with no benefit. Keeping inference to the handful of genuine judgment calls keeps per-run LLM cost flat as the workload scales.
- Reasoning quality. Each agent's context carries only what it needs to reason about, not the raw byproducts of every tool call.
- Testability & auditability. Deterministic tools are unit-testable; their outputs are reproducible and go straight into a structured JSON contract.
- Prevents scope creep. When a new requirement arrives ("add a second validation check"),
the answer is clear: add a
@tool, not a new agent. - Failure containment. Asking a deterministic function to interpret "there's a promotion next week" will fail; asking an LLM to compute a replenishment formula invites hallucinated arithmetic. Matching each task to the right executor removes both failure modes.
Worked application (13 components → 6 agent, 7 tool)¶
| Work | Deterministic? | Implementation |
|---|---|---|
| Parse natural-language request | No | agent reasoning |
| Load file from S3 | Yes | @tool |
| Select covariates from data quality + context | No | agent reasoning |
| Invoke Chronos2 endpoint | Yes | @tool |
| Interpret forecast anomalies in context | No | agent reasoning |
| Calculate order quantity from formula | Yes | @tool |
| Check order vs warehouse/budget constraints | Yes | @tool |
| Generate order rationale | No | agent reasoning |
| Generate chart / write JSON to S3 | Yes | @tool |
| Retry-with-adjusted-constraints vs escalate | No | agent reasoning |
| Summarize in business language | No | agent reasoning |
| Score forecast accuracy vs actuals | Yes | code-based evaluator |
The second boundary: in-process vs Gateway¶
The determinism test decides tool vs agent. A second, orthogonal test decides where a tool lives — it turns on trust, durability, and cost-of-mistake, not determinism:
"If the agent hallucinates and calls this tool wrongly, does the mistake propagate to external systems or stop at the agent's memory?"
- Stops at the agent → in-process
@tool. IAM role + type system already bound it; adding a gateway adds latency/cost with no safety gain. - Propagates externally → Gateway with authorization (JWT identity + Cedar policy) and an audit trail of "who asked for this write, and what was persisted."
In the canonical instance, exactly one of eight tools (save_decision, the authoritative
order write) needs the Gateway + Cedar policies;
generate_forecast_chart also writes S3 but stays in-process because a chart is a visualization,
not a decision of record. "That proportion is the norm, not the exception." Placing a high-value
deny policy on the write boundary (not on an upstream calculate step) is what makes it meaningful
— the agent could re-run a calculation until it passed, but the durable effect happens only at the
write.
Extends to evaluator design¶
The same logic applies one layer up, to which evaluator type to use (see two-layer evaluation): forecast accuracy is arithmetic with a correct numeric answer → code-based evaluator; agent behavior (did it flag the violation, was the rationale coherent) is subjective → LLM-as-a-judge. "Use the evaluator type that matches the question, not the one that feels more sophisticated."
Relationship to other patterns¶
- Complements patterns/specialized-agent-decomposition: decomposition decides how many agents; this pattern decides what inside each agent is an agent vs a tool.
- Pairs with patterns/tool-surface-minimization: fewer, well-scoped deterministic tools per agent.
- Produces the structured JSON contract that deterministic tool outputs flow into.
Seen in¶
- sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore — the
organizing thesis of the four-agent inventory pipeline; Bedrock inference happens in exactly the
four agents, everything else runs as deterministic
@tools, and onlysave_decisioncrosses the Gateway boundary. - sources/2026-09-28-redpanda-your-ai-kill-switch-is-in-the-wrong-place — applies the determinism discipline to authorization: a policy decision at the agent boundary must be a deterministic typed evaluation, not the agent's (or a guard-model's) judgment. "The same request in the same context must get the same decision every time, with nothing left to chance or the agent's own understanding of its boundaries." This is the security-boundary instance of "judgment belongs to the LLM, deterministic evaluation belongs outside it" — the allow/block/hold-for-human verdict is exactly the kind of reproducible, auditable decision that must not live inside the non-deterministic agent (Redpanda's OBPE).
Related¶
- patterns/specialized-agent-decomposition — how many agents vs what is a tool within one.
- patterns/tool-surface-minimization — keep each agent's deterministic tool set small.
- patterns/data-contract — deterministic tool outputs become the typed contract between agents.
- concepts/structured-output-reliability — why deterministic outputs are preferred at boundaries.
- concepts/llm-hallucination — the failure mode avoided by not asking an LLM to compute.