Suppose an AI workflow finds a supplier's website. The name looks right. The model returns MATCH with high confidence. The pipeline updates the supplier record.
One problem: the website belongs to a different company with the same name. That error now feeds enrichment, routing, and the next automated decision. A plausible interpretation has become an operational fact.
The architectural question is what had to pass before that write was allowed.
A prompt can ask a model to be careful. Software can require evidence, enforce permissions, and stop a transaction. Put those controls between the model's proposal and the business consequence.

For the supplier example, give the model a narrow job: interpret the supplied search results and propose a candidate. Require a candidate identifier, supporting source references, and one of three labels: MATCH, NO_MATCH, or UNKNOWN.
Those labels describe the model's judgment. Your application owns acceptance.
Reject malformed responses and identifiers that were never supplied. Then check the proposed company against trusted records or an approved alias map. A verified mismatch is a veto. An unresolved identity goes to review. The model's assertion that two companies are the same cannot also serve as the evidence that proves it.
| What arrives | What the system establishes | What happens |
|---|---|---|
MATCH |
Domain blocked by policy | Reject with reason |
MATCH |
Verified company mismatch | Reject with reason |
MATCH |
Identity evidence missing | Send to review |
UNKNOWN |
No mandatory veto applies | Send to review |
NO_MATCH |
No mandatory veto applies | Reject candidate |
MATCH |
Required checks pass; action authorized | Accept and record decision |
Apply mandatory vetoes first. Give each failure a recorded reason. That makes the workflow inspectable: an operator can see which condition permitted a result and which condition stopped it.
The same discipline applies to document answers. Code can verify that a source was retrieved, the requester has access, and a quoted passage exists. Establishing whether that passage supports the answer requires a separate evaluation. Another model can assist with that judgment; its assessment also carries uncertainty.
Structured output makes these checks easier to implement. It does not certify truth. OpenAI explicitly documents that schema-conforming outputs can still contain mistakes. Structured Outputs documentation.
The guarantee needs a precise boundary. Given the same proposal, evidence, policy version, and relevant system state, deterministic decision code should produce the same decision. A fresh model call can change the proposal. The final outcome can change with it.
Five confidence scores above an acceptance threshold demonstrate agreement on those runs. They cannot establish repeatability on the next run. Before a confidence score drives a threshold, calibrate it against labeled examples from the actual task.
Treat outcome stability as something to measure. Replay saved proposals to test the decision code. Rerun the model on a fixed evaluation set to measure how often the workflow changes its answer. Include ambiguous names, missing evidence, contradictory sources, and cases close to the acceptance threshold.
Several engineering choices reduce unnecessary variation and contain the failures that remain:
- Normalize before interpreting. Standardize whitespace, resolve approved aliases, and deduplicate candidates. Strip URL parameters only when they are known to be irrelevant. Preserve distinctions that change company identity or the underlying resource.
- Version the whole decision path. Record the model, prompt, schema, retrieval configuration, source snapshot, and policy code. Keep context ordering and supported generation settings fixed. Provider-side variation can remain even with fixed settings and seeds. OpenAI reproducibility guidance.
- Make exceptions explicit. Give
UNKNOWNa named reviewer, supporting evidence, and a response target. Allow bounded retries for recoverable formatting or service failures. Retrying a policy rejection until the model supplies an acceptable answer defeats the policy. - Protect execution. Resolve equal rankings with a stable identifier. Keep mandatory rules outside model control, including when retrieved text contains instructions. Recheck authorization before a write. Enforce a unique operation key in the destination transaction so repeated attempts cannot create duplicate records.
In a Dataiku or Snowflake implementation, separate proposed records from accepted records. Give the model-facing stage permission to create proposals. Reserve accepted-record writes for the component that enforces the checks. Preserve the original proposal and permitted evidence alongside the final decision so an investigation can reconstruct what happened.
This design has a cost. Stricter gates send more cases to review. Review queues need owners and capacity. Measure incorrect automatic acceptances, review volume, resolution time, and cost per accepted result. Choose the operating point deliberately. A system that escalates everything has automated very little.
Start with the most expensive mistake one workflow can make. Identify the evidence that would expose it, encode the conditions that must stop it, and test whether any path reaches the business action without passing those checks.
Which action can your model trigger today without independently verified evidence?