AI Ops: How to Run AI Systems in Production Without Surprises

AI Ops for LLM apps: tracing, evals in CI, model version pinning, drift detection, cost and latency SLOs, and incident response for AI systems in production.

S
Softzee EngineeringSeptember 15, 2026 · 6 min read

Traditional software fails loudly: an exception, a 500 error, a crashed pod. AI features fail quietly. The service returns 200, the latency looks normal, and the answer is wrong, off-policy or ten times more expensive than yesterday. AI Ops is the practice of running LLM-based systems so you notice those failures before your customers do.

Most of what you already know from DevOps still applies: monitoring, CI, on-call, change control. What changes is what you monitor and how you test. This guide covers the operational pieces we put in place before an AI system goes live.

Observability for LLM apps: trace every request end to end

A single user request to an AI feature can involve a query rewrite, a retrieval step, a reranker, two or three model calls, several tool calls and a guardrail check. When the final answer is wrong, you need to see which step went wrong. Request-level logs are not enough. You need traces.

Each trace should capture, per step:

  • The full prompt sent (or a reference to it) and the full response
  • Model name and exact version, plus parameters such as temperature and max tokens
  • Input, cached and output token counts, and the calculated cost
  • Latency, including time to first token for streamed responses
  • Retrieved document IDs and their relevance scores
  • Tool calls with arguments, results and errors
  • User, tenant, feature and prompt template version

OpenTelemetry has emerging semantic conventions for generative AI spans, and most LLM observability platforms, both open source and commercial, can ingest them. Using a standard means your AI traces can sit alongside your existing application traces instead of in a separate silo.

Two practical warnings. First, prompts and responses often contain personal data, so apply the same retention, masking and access rules you would apply to any customer data. In Saudi Arabia that means PDPL obligations apply to your logs too. Second, sample intelligently: keep every trace for errors, low user ratings and high-cost requests, and a percentage of the rest.

Evals in CI: tests for behavior, not just code

A one-line prompt change can fix one case and break five others. Without automated evaluation, nobody will notice until users complain. Treat prompts, model choices and retrieval settings as code, and gate them with evals in your CI pipeline.

Build the eval set from reality

Start with 50 to 200 representative cases taken from real traffic, plus known edge cases and past failures. Each case has an input and a way to judge the output: an exact expected value, a schema check, a set of facts that must appear, or a rubric for an LLM judge.

Mix cheap and expensive checks

Deterministic checks (valid JSON, required fields present, no forbidden phrases, correct tool called) are fast and reliable. Run them on every commit. LLM-as-judge scoring for tone, helpfulness and faithfulness is slower and noisier, so run it on pull requests that touch prompts or models, and calibrate the judge against human ratings first.

# .github/workflows/ai-evals.yml (excerpt)
on:
  pull_request:
    paths: ["prompts/**", "config/models.yaml", "rag/**"]
jobs:
  evals:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: python evals/run.py --suite core --min-pass-rate 0.95
      - run: python evals/run.py --suite safety --min-pass-rate 1.0

Notice the safety suite requires a perfect score. Some failures, such as leaking another tenant's data or agreeing to an unauthorized refund, are not acceptable at any rate. Our QA automation work increasingly includes these behavior suites alongside regular tests.

Model version pinning and controlled upgrades

Providers update models, retire old versions and sometimes change behavior behind an alias. If your code calls a floating alias like "latest", your product can change on a Tuesday without anyone on your team deploying anything.

  • Pin exact model versions in configuration, not in scattered code.
  • Track provider deprecation schedules and put retirement dates in your team calendar well ahead of time.
  • Treat a model upgrade as a release: run the full eval suite, compare cost and latency, then roll out gradually with a canary or a percentage split.
  • Keep the previous version available as a fast rollback path until the new one has proven itself in production.
  • Version prompt templates alongside the model, since a prompt tuned for one model often behaves differently on another.

The same discipline applies to embedding models. Changing the embedding model means re-indexing your entire corpus, so plan it as a migration, not a config tweak.

Drift: when nothing changed but results did

Even with pinned versions, quality can degrade over time. The causes are usually on your side: users start asking about a new product, the knowledge base goes stale, a source system changes its data format, or a seasonal shift changes the mix of questions. In Gulf markets, for example, traffic patterns and topics can shift noticeably around Ramadan and major sales periods.

Watch for drift with signals you can compute continuously:

  • Distribution of detected intents and topics over time
  • Rate of "I don't know" or fallback responses
  • Escalations to human agents and user thumbs-down ratings
  • Retrieval relevance scores trending down
  • Average tokens per request and tool calls per task trending up

Run a sample of live traffic through your automated judges daily, so you have a production quality score and not just an offline one. A falling score with no deploy is the signature of drift.

Close the loop with human review

Automated signals tell you something moved. They rarely tell you why. Set up a weekly review where someone who knows the domain reads a sample of real conversations, especially the flagged ones: low ratings, escalations, refusals and unusually long or expensive sessions. Tag each failure with a cause (retrieval miss, outdated content, wrong tool, prompt gap, model limitation) and track those counts over time. This review is usually where you find the fixes with the biggest payoff, and it feeds new cases straight into your eval set. Thirty minutes a week from a support lead and an engineer is often enough to catch problems that dashboards miss.

Cost and latency SLOs

For AI features, service level objectives need more than availability and error rate. Define targets your product and finance teams both agree on:

SLOExample targetWhy
Time to first tokenp95 under 1.5 seconds for chatPerceived speed in streaming interfaces
End-to-end latencyp95 under 800 ms per turn for a voice agentVoice conversations break down with long pauses
Cost per successful taskBelow a set amount per resolved ticketProtects margins on fixed-price plans
Quality scoreJudge pass rate above 90% on sampled trafficCatches silent regressions

The numbers above are examples. Set yours from your own baseline and product needs. Alert on burn rate, not on single spikes, and put hard per-tenant and per-task spending caps in code so a runaway agent loop cannot generate an unbounded bill. Plan for provider limits too: rate limit errors and regional outages happen, so a tested fallback to a second model or provider is worth having for critical flows.

Incident response for AI systems

AI incidents look different. A model gives harmful advice, an agent takes an action it should not have, a prompt injection extracts a system prompt, or costs spike tenfold overnight. Your runbooks should cover these explicitly.

  1. Kill switches. Be able to disable a feature, a tool or an agent action instantly through a flag, without a deploy.
  2. Safe fallback. Decide in advance what users see when the AI path is off: a human handoff, a static answer, a form.
  3. Triage from traces. Find every affected request by prompt version, model version or tool, and assess scope.
  4. Fix and add a test. Every incident becomes a new eval case, so the same failure cannot ship again.
  5. Review. Run a blameless post-incident review, as you would for any outage.

None of this requires a new team. It requires your existing DevOps practice to treat prompts, models and retrieval indexes as production components with owners, tests and rollback plans.

How Softzee can help

We build and operate production AI systems, including voice agents and support assistants, with tracing, evals and cost controls built in. If you have AI features in production and limited visibility into how they behave, book a call and we will help you put the right AI Ops foundations in place.

AI OpsLLM observabilityEvalsDevOpsMLOps

Have a project in mind?

Tell us what you are trying to build. You will get an honest take on scope, timeline and cost, usually within one business day.

Keep reading

All articles