A system isn't in production because it runs
It is in production when a real person's work depends on its accuracy—and a real stakeholder gets the call when it fails.
A model demo answers a narrow question: can this input produce a convincing output? A production system has to answer harder questions. What happens when the input is malformed? When the provider times out? When a response arrives late and overwrites a newer one? When a field is plausible but wrong? When an auditor asks which evidence supported the decision?
Those questions change the architecture. In a financial document-processing workflow, I start by defining the document classes, field-level risk, business owner, latency budget, data boundary, and escalation path. Only then does model selection become useful.
Evaluate the work, not the brand
Model routing should follow measured fitness for the task. I route by document class, field risk, confidence threshold, latency budget, and evaluation performance—not by the reputation of a provider or a generic benchmark. Golden test fixtures and field-level scoring make regressions visible. Segmented failure analysis tells us whether the problem is retrieval, normalization, model behavior, or data quality.
Make failure an explicit state
Reliable agentic systems need persisted workflow state, idempotent tool calls, bounded retries with exponential backoff, and clear terminal outcomes. A timeout should not become an ambiguous blank. A late response should not silently replace current state. A partial result should not masquerade as a complete one.
Correlation IDs connect a user action to retrieval, model, tool, and integration events. Latency, error, and cost telemetry reveal where the system is drifting. Versioned prompts, models, and sources make an output reconstructable.
Put people where judgment belongs
Human review is not a universal fallback. It should be designed around risk: high-impact actions, low-confidence evidence, policy exceptions, and cases where responsibility cannot be delegated to a probability. The system should explain why a case was escalated and give the reviewer the evidence needed to decide.
Close the loop
Production feedback becomes useful when it is classified. I separate workflow gaps, retrieval failures, model behavior, user-experience friction, data-quality issues, missing integrations, and examples that belong in the evaluation suite. That classification lets Product and Engineering change the right layer.
The practical definition of production is not “the model answered.” It is “the system can be trusted, observed, corrected, and owned.” That is the work around the model call—and it is usually where the product earns its value.