Spectre
// PUBLISHED03.10.26
// TIME9 MINS
// TAGS
#AI#LLM#RELIABILITY#PRODUCTION
// AUTHOR
Spectre Command

Y

our regular software either works or it doesn't. A null pointer exception is deterministic. Run the same inputs, get the same crash, fix the bug, done. LLM features fail differently and if your team is building them the same way they build everything else, you're going to find out at the worst possible time.

Reliable ai feature production isn't about making AI perfect. It's about building a system that degrades gracefully when the AI is wrong, slow, or unavailable. The engineering discipline is different. The questions you need to ask your team are different. And most teams even good ones haven't asked them yet.

This is what I'd want to know as a CTO before shipping any AI feature to production.

Hallucination is not a bug you can fix. It's a property you have to engineer around.

The biggest mistake I see CTOs make: treating hallucination as a problem the model will eventually solve. Waiting for the next version. Assuming the fine-tuned model handles it. It doesn't. Not completely. Probably not ever for all inputs.

Hallucination the model generating confident, plausible, wrong output is a fundamental characteristic of how these models work. You don't eliminate it. You contain it.

The question to ask your team: what happens when the model is wrong? Walk me through it. If the answer is "the user sees wrong information," that's a reliability problem that sits at the product level, not the model level.

Containment strategies depend on the use case. For factual retrieval, pair the model with a retrieval system [→ RAG vs Fine-Tuning vs Prompt Engineering] covers this in detail so the model is synthesising from a bounded set of verified documents rather than drawing from training data. For any feature where a wrong answer has real consequences (medical, legal, financial, customer-facing commitments), human review in the loop isn't optional. For lower-stakes features, output validation structured output formats, schema enforcement, confidence thresholds catches a large fraction of obvious failures before they reach the user.

The gotcha: teams build output validation for the happy path. They test it on examples that work. The hallucination problem is that the model is most confidently wrong on exactly the inputs that look like normal inputs. Your validation needs to handle the case where the model produces something structurally valid but semantically wrong.

Latency is a first-class feature requirement. Treat it like one.

LLM calls are slow. P50 latency for a frontier model generation might be 2–4 seconds. P95 might be 8–12 seconds. P99 on a bad day in a degraded region could be 30+ seconds. These numbers are not hypothetical they're what production telemetry shows for teams who measure.

Most teams don't set explicit latency budgets for AI features. They ship, they watch, they tolerate whatever the model returns. The problem is that users don't tolerate it. On Indonesian mobile networks especially, where connectivity drops and resumes unpredictably, a 15-second spinner is a feature that feels broken.

Ask your team three questions. First: what's the p95 latency of every LLM call in the product, measured from our servers? Not the model's reported latency. Our end-to-end number. Second: what's the timeout? Third: what happens when we hit that timeout?

Streaming is the primary mitigation for perceived latency. Start rendering output token-by-token as the model generates it. Users tolerate 10 seconds of watching text appear far better than 10 seconds of a blank loading state. It doesn't make the model faster, but it makes waiting bearable. If your team hasn't implemented streaming on user-facing generation features, that's the first thing to fix.

The timeout question matters more than most teams realise. If you have no timeout and I've seen production systems with none a single slow model call can hold a connection open for 60+ seconds, consuming server resources and blocking retries. Set explicit timeouts. 10–15 seconds for most cases. Shorter for features where speed is critical. Then engineer the fallback.

The fallback your team hasn't built

Every LLM-dependent feature needs a defined behaviour for when the LLM is unavailable, too slow, or returns something unparseable. Most teams haven't defined this. Some haven't thought about it.

The options, roughly in order of engineering effort: fail gracefully with a user-visible message ("this feature is temporarily unavailable"), return a cached or default response, fall back to a simpler non-LLM implementation, or retry with exponential backoff.

Which fallback makes sense depends on the feature. A document summarisation feature that returns "summary unavailable view the full document" is fine. A real-time support chatbot that silently drops messages is not. The distinction matters and it needs to be a conscious product decision, not whatever happens when the catch block is empty.

The circuit breaker pattern applies here directly. [→ API design for high-throughput systems] covers the general principle if the LLM provider starts failing, stop hammering it and activate the fallback path instead of queuing thousands of requests that will all time out. In practice this means tracking error rates per model endpoint and tripping a circuit when they exceed a threshold.

One thing teams consistently underestimate: the LLM providers themselves have incidents. OpenAI, Anthropic, Google all of them have had regional outages and degraded service windows in the past 18 months. If your product has no fallback and the provider goes down, your AI feature goes down with it. That's fine if it's a nice-to-have. It's not fine if it's core to your product's value proposition.

What "observability" means for AI features specifically

For regular software, observability means logs, metrics, and traces. You know what I mean: did the request succeed, how long did it take, where did it fail. The OpenTelemetry tooling handles most of this.

AI features need that layer plus something else. You need to know whether the AI is doing a good job, not just whether it's running. Those are different questions.

The minimum instrumentation: log the input, the output, and the model used for every LLM call in production. Store it somewhere you can query. This sounds obvious and most teams don't do it, because they're moving fast and "we'll add proper logging later." Later never comes, and when a user reports that the AI said something wrong, you have no record of what it said or why.

From that log data, you can start building quality metrics specific to your feature. For a summarisation feature: are outputs within expected length ranges? Are they in the right language? For a classification feature: are confidence scores distributed the way you'd expect, or is the model increasingly uncertain about something that used to be easy? For a chatbot: what's the rate of follow-up "I don't understand your answer" messages?

These are proxy metrics for quality. They're not perfect. But they're the difference between "we think the AI is working fine" and "we have data that suggests the AI is degrading on this input type since the model version changed."

Real-world example: an Indonesian fintech chatbot goes wrong silently

A fintech startup running a customer support chatbot for loan inquiries had a good product and solid initial quality. Six months after launch, a support team lead flagged that more customers were following up the chatbot with a call to human support. The rate had roughly doubled.

The team had no LLM-specific observability. No output logging. They knew the system was up and handling requests. They didn't know what it was saying.

When we instrumented logging and pulled a week of outputs, the pattern was clear. A model provider update had changed how the model handled edge cases in Bahasa Indonesia mixed with financial terminology. The chatbot was answering questions correctly in standard Indonesian but falling apart on the code-switching that real users actually write "berapa bunga kalau aku ambil tenor 12 bulan buat yang KTA type B?"

None of their existing monitoring caught it. The feature was technically up. The quality had quietly degraded for weeks.

They added input/output logging, language detection on outputs, and a human review sample on 2% of conversations. The degradation was caught within 48 hours of the next model update. The escalation-to-human rate returned to baseline within a month.

FAQ

Q: How do I know if hallucination is a real risk for our specific use case?

A: Ask what the consequences are when the model is wrong. If the answer is "a user sees an incorrect sentence and moves on," the risk is low. If the answer is "a user acts on incorrect financial, medical, or legal information," or "a user is quoted an incorrect price," the risk is significant and needs active mitigation retrieval grounding, human review, structured output validation, or some combination.

Q: What's a reasonable latency SLA for an LLM feature?

A: It depends entirely on the UX context. A background document processing job has different expectations than a real-time chat interface. For user-facing generation, anything over 3–4 seconds without streaming starts damaging perceived quality. With streaming, users tolerate 8–10 seconds noticeably better. Set a hard timeout at 15 seconds for most cases and design your fallback around it.

Q: Should we build multi-provider failover between OpenAI and Anthropic?

A: It's worth the investment if the AI feature is critical to your product's value proposition. The engineering cost is real different APIs, different prompt formats, different output behaviours to account for. An LLM gateway layer that abstracts the provider makes this manageable. If the feature is peripheral, a graceful degradation fallback is usually enough.

Q: How do we test AI features before shipping?

A: Build an eval set from real production inputs, anonymised. At minimum 50–100 examples covering edge cases, not just the happy path. Define what "correct" looks like for each example even rough human labels are better than none. Run the eval on every significant prompt change or model version change before it goes to production. This is unglamorous work and it's what separates teams who catch regressions from teams who get user complaints.

Q: My team says the model is "good enough." How do I validate that?

A: Ask them to show you the eval set and the results. If there's no eval set, "good enough" is an opinion, not a measurement. The absence of a structured evaluation process is itself a signal that quality is being managed reactively rather than proactively.


AI features don't fail loudly. They fail quietly, gradually, in ways your existing monitoring wasn't designed to detect. The engineering discipline required is different from what your team already knows not harder, but different. Getting it right means asking these questions before launch, not after a user reports something wrong. If you're evaluating whether your team has the foundations in place, the [→ AI-Native Startup Architecture guide] is a good place to start.

External Documentation:

  • [OpenTelemetry Docs] The standard observability framework for distributed systems, applicable to LLM tracing pipelines.
  • [Anthropic Model Cards] Capability and limitation documentation per model; useful for setting realistic reliability expectations.
// END_OF_LOGSPECTRE_SYSTEMS_V1

Is your current architecture slowing you down?

Stop guessing where the bottlenecks are. We partner with founders and CTOs to audit technical debt and execute zero-downtime system rewrites.

Book an Architecture Audit