Each stage in this flow has a distinct job. Click any box in the diagram to see what it does — we'll go deep on each one over the next sections.
A raw request — a message, a ticket, an API call — with no guarantee of clean phrasing or intent.
An agent assignment, plus a confidence value the rest of the system can act on.
Does one job well, not many jobs adequately.
Knows what it doesn't do.
Only the tools and data it needs.
Knows when it's unsure.
Reliable, predictable shape.
The agent's confidence score falls below threshold — it's unsure this is even the right category.
Typical thresholds — most systems commit above ~70–85% confidence and escalate below it. A middle "gray zone" (~50–70%) often routes to a more capable agent first. Exact numbers are tuned per system.
Not "I'm unsure" — "this genuinely isn't mine." A billing agent asked to debug an API should hand off immediately.
Legal threats, large sums, fraud signals — some topics get human review regardless of confidence. A deliberate design choice.
The user keeps saying "that's not it." After N attempts, stop looping and hand off instead of forcing the automated path.
A second copy of the same agent or service running elsewhere. Same capability, different instance.
In practice — Netflix runs services across multiple AWS zones; if one has an outage, traffic shifts to an identical instance elsewhere and most viewers never notice.
Primary model API down → route to a secondary provider. Usually a downgrade, but keeps the system running.
In practice — a support platform's primary LLM provider has an outage, so new requests route to a backup model from a different vendor — slightly different answers, but the queue keeps moving.
A smaller, faster, less capable version that gives an answer rather than timing out repeatedly.
In practice — a coding assistant's deep multi-file agent keeps timing out on a huge repo, so it falls back to a lighter single-file agent that at least answers instead of leaving the user with nothing.
For read-heavy tasks, a clearly-labeled stale answer can beat returning an error.
In practice — a weather app's live-conditions API fails to respond in time, so it shows last hour's reading labeled "as of 40 min ago" rather than a blank screen.
One more layer worth knowing: a circuit breaker stops calling a dependency entirely for a cooldown period after repeated failures.
In practice — Netflix's Hystrix pattern: if a service fails repeatedly, the app stops calling it for 30 seconds and shows a fallback (e.g. popular titles) instead. One test request after the cooldown decides whether normal traffic resumes.
Borderline calls involving values or risk — not something the system should decide alone.
For high-stakes actions, a human simply approves or stops the agent's proposed action before it executes.
The request doesn't fit any category — someone has to actually solve it, script or no script.
Every resolved escalation is training signal — for tightening scope, thresholds, and routing.
The best systems treat human escalation as a precision instrument, reserved for what truly needs it — not a safety blanket for cases the system wasn't designed carefully enough to handle.
Right side: the agent doesn't know what to do — escalate to a person. Left side: something broke technically — retry, then fail over. Keep these signals separate; a spike in one means "improve the agents," a spike in the other means "fix the infrastructure."