Skip to content
Malik Hamza Shabbir
Agents in Productionllm-fallbackfail-closedai-guardrailsagents-in-production

Our LLM Fallback Worked Perfectly. That Was the Problem.

HSMalik Hamza Shabbir11 min read

In short

A fallback that quietly returns something during a model outage is the riskiest path in any AI feature that publishes under someone else's name: ours posted generic English replies to Swedish businesses' Google profiles for days while every dashboard stayed green. My rule now is that every degraded path refuses instead of publishing, through a separate publish gate that checks language and boilerplate and returns HOLD, with an alert whenever a fallback fires. Fail open is still right for some paths, so decide per path by who gets hurt.

Isometric pipeline of glowing message blocks passing through a gate, with one block held in a side tray under a warning beacon.
On this page

Short answer: a fallback that quietly returns something during a model outage is the most dangerous path in any AI feature that publishes under someone else's name, because it turns an outage you would notice into wrong content you won't. On a review-reply product I lead, a fallback path posted generic English boilerplate to live Google Business profiles of Swedish businesses for days, and every dashboard we had stayed green. The fix was to make every degraded path refuse instead of publish: check the reply's language, reject boilerplate, hold the reply when the primary model is down, alert when a fallback fires, and make publishing its own gated step. Fail open is still right on some paths, so I decide per path by asking who gets hurt if the system guesses.

What actually happened when our fallback kicked in?

I lead the build of a review-reply product for businesses in Sweden. It answers their Google reviews for them, in the business's voice and in the customer's language, and publishes the reply on the business's Google Business Profile. The business's name sits above it. Their customers read it, and so does anyone who looks them up later, because a reply stays under the review until somebody removes it.

During an outage of our primary model, a fallback path took over. It didn't error. It produced generic English boilerplate, the kind of thanks-for-your-feedback reply that fits any review and therefore fits none, and posted it to live profiles of Swedish businesses.

It went on for days before anyone noticed.

From the system's point of view, nothing went wrong. Requests returned OK. Replies were posted. The response-rate numbers, the ones this kind of product is judged on, looked great. If you had opened any of our dashboards during those days, you would have seen a healthy system doing its job.

This is the bit that still bothers me. We didn't have a monitoring gap in the usual sense. We had monitoring that measured the wrong thing very accurately.

Why did every dashboard stay green?

Because every dashboard counted whether something happened, and a fallback exists to make sure something always happens.

A normal outage announces itself: errors spike and someone gets paged. A fallback is built to prevent exactly that, and ours did. Nothing errored. Keeping the queue moving is the fallback's whole job, so there's no pile-up to make anyone suspicious either. A fluent wrong reply doesn't trip anything that measures throughput.

This isn't a small-team problem. In September 2025 Anthropic published a postmortem of three overlapping infrastructure bugs ↗ that degraded Claude's output quality while requests kept succeeding. At its worst the degradation hit 16% of Sonnet 4 requests, and one symptom was Thai characters turning up in English replies. Anthropic says its own evaluations didn't capture what users were reporting.

A 200 tells you the pipe is open. It tells you nothing about what came through it. If your AI publishes, you need at least one signal that looks at the output itself, even if it's a person reading a handful of live replies every morning. If you trace your model calls, put the model that answered and a fallback flag on the span, and alert on them.

Why do LLM fallbacks fail like this?

Amazon worked this out long before LLMs. The Builders' Library essay on avoiding fallback in distributed systems ↗ puts it plainly: "Fallback paths rarely get exercised, and thus they tend to hide latent bugs." Code that only runs during an outage is the code you've tested least, running at the worst moment. Amazon's answer is to make the primary more reliable instead.

The circuit breaker pattern makes it worse for text. Microsoft's circuit breaker guidance ↗ allows the open state to hand back a default value instead of calling the failing service. For a stock widget, a cached price is a sane default. For words that get published under a business's name, the default is the incident. The same page tells you to raise an alert when the breaker changes state, which is the half most of us skip.

Gateways make it easy. In LiteLLM's reliability config ↗, a fallback is one line. When it fires, the evidence is in the logs, in spend-log fields like attempted_fallbacks and original_model_group, and in an x-litellm-attempted-fallbacks response header. Nothing checks whether the fallback's output is any good, and a gateway can't: it has no idea what a good review reply looks like. That leaves a fallback free to run for days with a response header as its only trace.

Then there's language. Cohere's research on language confusion in LLMs ↗ (EMNLP 2024) found that even the strongest models they tested didn't reliably answer in the user's language, and that it got worse with complex prompts and higher temperature. A long English prompt wrapped around a two-line Swedish review is that setup exactly. The language check I added because of the fallback ends up protecting the primary model too.

What does fail closed mean for an AI that publishes?

The rule I apply to everything that posts publicly now: every degraded path refuses rather than publishes. If a reply can't be produced the proper way, it isn't produced. It waits.

Five changes came out of it:

  • Language match. Detect the language of the review and the language of the reply before anything is published. If they differ, hold.
  • Boilerplate detection. Compare the reply against known templates and generic phrasing. If it's too close, hold. A generic reply isn't a neutral outcome either: in BrightLocal's 2026 consumer survey ↗, half of consumers said they view templated or generic review replies negatively.
  • No substitution. If the primary model is unavailable, the reply waits in a queue as a draft. Nothing stands in for it.
  • Loud fallbacks. A fallback firing raises an alert that reaches a person. A log line doesn't count.
  • Publishing is its own step. It has its own gate and runs separately from generation, instead of being the last line of the generate function.

That last one is what makes the others hold. When publishing is the tail end of generation, every check lives inside the same code as the fallback, which is the code that never gets exercised. When publishing is separate, the gate doesn't care how the draft was made. It looks at the draft.

TEXT
Before (simplified):
  new review -> generate (primary, else fallback) -> post to Google -> OK -> response rate +1

After:
  new review -> generate (primary only) -> save draft
  draft      -> publish gate -> PUBLISH -> post to Google
                             -> HOLD    -> queue + alert -> retry later, or a human decides

A reply that turns up tomorrow costs the business very little. A wrong one sits on their profile until someone removes it. It's the same instinct as a docs chatbot that refuses low-confidence answers instead of guessing ↗: silence is a valid output, and sometimes the best one on offer.

What does a publish gate look like in code?

Illustrative, not our production code, but it's the shape I'd start from.

PYTHON
# Illustrative. Thresholds are placeholders; tune them on your own replies.
MIN_LANG_CONFIDENCE = 0.80
MAX_TEMPLATE_SIMILARITY = 0.85

def publish_gate(d) -> tuple[str, str]:
    # d.model_used comes from the gateway response, not from a flag our code sets.
    # d.allowed_fallbacks is empty by default; a person adds to it, per account.
    if d.model_used not in (d.primary_model, *d.allowed_fallbacks):
        return hold(f"answered by {d.model_used}, which may not publish")

    review = detect_language(d.review_text)    # deterministic detector, no LLM call
    expected = review.code if review.confidence >= MIN_LANG_CONFIDENCE else d.business_language
    reply = detect_language(d.reply_text)
    if reply.code != expected:
        return hold(f"language mismatch: expected {expected}, got {reply.code}")

    similarity = max_similarity(normalise(d.reply_text), KNOWN_TEMPLATES)
    if similarity > MAX_TEMPLATE_SIMILARITY:
        return hold(f"reply looks like a template ({similarity:.2f})")

    return ("PUBLISH", "")

def hold(reason: str) -> tuple[str, str]:
    alert(reason)                  # reaches a person; a log line is not enough
    return ("HOLD", reason)        # never a substitute reply

A few choices are deliberate. The gate can only return PUBLISH or HOLD, so no path invents a reply. The fallback check reads which model answered from the gateway's response, not from a flag the generation code was meant to set. Star-only reviews have nothing to detect, so the expected language comes from the business rather than the check being skipped. And hold always alerts, because a HOLD nobody hears about is a silent backlog, which is the original problem moving slower.

OpenAI's agent guidance lands in the same place. Its guardrails and human review docs ↗ say to pause before side effects and wait for human approval, and in the Agents SDK ↗ an output guardrail that trips raises OutputGuardrailTripwireTriggered and stops the run. Posting a review reply is a side effect. So is sending an email.

When is fail open the right call?

I don't think fail closed is a universal virtue, and my own counterexample comes from a different system.

On an outbound email engine I work on, addresses go through an email verification provider before anything is sent. When that provider's credit balance ran out, verification started returning "unknown". The engine treated unknown as bad. It stopped sending without telling anyone, and then went further and blocklisted perfectly good domains. That was fail closed by the book, and it was wrong. I changed it to fail open on a first unknown and to raise a visible alert on the system page when the balance is empty.

Side by side, the two incidents differ in one way that matters: who gets hurt when the system guesses.

A first unknown address that goes out anyway risks a bounce. That cost is small and it lands in the sender's own numbers. Blocking cost the client their sending, and nobody knew why. A substituted AI reply sits at the other end. It's public and it carries a business's name. It stays up until someone spots it.

There's a dev.to post arguing that a reviewer stage should never block a decision ↗. For an internal, advisory reviewer, maybe. For anything published under someone else's name I'd turn it round: the reviewer is the one thing that should be allowed to say yes.

That's why there's no global policy in my systems any more. Each path gets decided on its own:

TEXT
For every path that can degrade:

Who sees the output?
  only us, internal tooling            -> fail open is usually fine
  a customer or the public             -> keep going

Whose name is on it?
  ours                                 -> judgement call, write it down
  a customer's, or their business's    -> fail closed

Can it be taken back before anyone sees it?
  yes (draft, queue, internal state)   -> fail open, with an alert
  no  (posted, sent, published)        -> fail closed: HOLD + alert

Every branch: a fallback firing is an alert, never only a log line.

The one thing both fixes share is the alert. The direction you fail should change from path to path. How loudly you fail shouldn't. The quiet part is what cost us days.

How do you test a guard that only runs when something breaks?

The same failure shape turned up on that outbound engine. The model once wrote an opening line naming the prospect's own product as the competitor that was beating them. Fluent, and addressed to a real person. There's now a check for that, and it blocks the send.

On the same engine I built an autonomous multi-turn reply agent. It ships switched off, behind a mode setting, and before every single send it re-runs one full stack of gates. Not the gates that seem relevant to that turn. All of them, every time.

The bug class I keep finding is guards that look implemented but can never actually run. The ways a guard ends up unreachable are boring, in any codebase: a check reads a field nothing ever sets, a guard sits below an early return, or a refactor moves sending into a new function and leaves the guard behind on the old one. Every one of those survives a read-through, because the guard is right there on the screen, correctly written, and never executed.

That's why I don't verify guards by reading the code. I verify them with tests that try to trip them: push a draft that must be blocked through the real send path and assert that nothing leaves. If a test can't make the guard fire, treat the guard as missing. They belong in CI next to your evals; I've written about scoring agent runs in CI ↗.

A rule that lives only in a prompt isn't a guard either. Replit's agent deleted a production database during a declared code freeze ↗ in July 2025. The freeze was an instruction to the model, and nothing in the system enforced it.

Who is liable when your AI publishes the wrong thing?

You are, or your customer is, and then they call you.

Air Canada argued that its chatbot was a separate entity. The tribunal in Moffatt v. Air Canada ↗ rejected that in February 2024 and held the airline liable for the bot's wrong bereavement-fare advice. The damages were C$650.88. The money was trivial. The precedent wasn't.

Cursor's support bot, "Sam", invented a one-device policy ↗ in April 2025, which led to cancellations and a public apology. Apple paused AI notification summaries for news apps ↗ in January 2025 after they produced false headlines. OWASP's 2025 Top 10 for LLM applications lists improper output handling (LLM05) and misinformation (LLM09) ↗ among the top risks, which is the standards-body way of saying don't pass model output straight to the world.

None of those needed an outage. They needed output that reached people with no check in between. Ours was quieter than any of them, with one extra twist: the businesses whose profiles carried those English replies never wrote them, and their customers had no way of knowing that.

What should you ask before your AI is allowed to publish?

The questions I ask before anything is allowed to post, send or publish on someone's behalf:

  1. What does each fallback path produce, word for word, and when did anyone last read it?
  2. If the primary model is down for a whole day, what does the customer's audience see? The right answer is "nothing new yet".
  3. Is publishing a separate step with its own gate, or the last line of the generation function?
  4. Does the gate know which model actually answered, from the gateway's response rather than a flag?
  5. Does it compare output language to input language, and catch near-template replies?
  6. Where do held items go, who sees them, and how does one get released?
  7. Does a fallback firing page a person, or write a log line?
  8. Is there a test that tries to trip each guard through the real publish path?
  9. If every request succeeded and every output was wrong, which number on which dashboard would move?

The last question is the one that would have saved us. For us, during those days, the honest answer was none of them.

If your product posts, sends or publishes under your customers' names and you're not sure what your fallbacks do on a bad day, that's a review I'm glad to do. The review-reply side lives on my reputation automation ↗ page, and the alerting half is part of SaaS security and monitoring ↗. Or get in touch ↗. I would rather talk you out of a clever fallback than build you one.

FAQ

What does fail closed mean for an AI feature that publishes content?

Fail closed means that when any part of the pipeline is degraded, the system publishes nothing rather than a substitute. The reply or post is saved as a draft in a held queue, an alert reaches a person, and it goes out only after it passes the normal checks. A late reply costs very little, while a wrong one published under a customer's name stays public until someone notices it.

Is it safe to use a fallback model when your main LLM is down?

Only if the fallback's output goes through the same publish gate as the primary's and every fallback firing raises an alert. The dangerous version returns a default string or an unchecked answer straight to users, because every request still succeeds and normal dashboards stay green. For anything that speaks under a customer's name, keep fallbacks off by default and hold the output instead.

How do I stop an LLM from replying in the wrong language?

Don't rely on the prompt alone. Detect the language of the input and of the output with a deterministic language detector, compare the two before publishing, and hold the reply if they differ. For very short or star-only inputs with nothing to detect, compare against the account's own language instead of skipping the check.

When should a system fail open instead of fail closed?

Fail open when a wrong guess is cheap and lands on you, such as letting an address with a single unknown email verification result go out. Fail closed when the output is public or carries someone else's name, because it cannot be taken back once people have read it. Decide path by path based on who gets hurt, and raise an alert either way.

How do I monitor an LLM fallback in production?

Record which model actually answered each request, taken from your gateway's response rather than a flag your own code sets, and alert whenever it isn't the primary. Add at least one signal that looks at the output itself, such as a language check or a person reading a sample of live replies, because error rates and success counts stay green while a fallback is quietly working. Then write tests that force the fallback path and assert that the publish gate holds its output.

Working on something like this?

I build web apps, AI features, and mobile products for clients. If this article matches a problem you have, tell me about it.

Start a conversation
HS

Malik Hamza Shabbir · Full-Stack & AI Engineer

I build full-stack and AI products solo: a reputation SaaS in production, RAG pipelines, and React Native apps. I write from what I ship, not from documentation summaries.

Run it on yours

Every number above came from code you can read.

The harness is open and the per-probe log ships with it. If you build one of these systems, or you are choosing between them with real money, the same method is available to you.

Send me an endpoint and I will write an adapter, run it on the same workload under the same rules, and publish the result. It goes in the table whether it wins or not.

Sponsored inclusion is disclosed in the article and the result publishes either way. Paying moves you up the queue, never up the table.

Related articles