Solution Architecture for AI Products: The Orchestration Layer

A reference solution architecture for AI products: gateway, orchestration layer, tools, memory, retrieval, guardrails and observability, plus what to build or buy.

S
Softzee EngineeringSeptember 26, 2026 · 7 min read

Most AI products start as a prompt wired straight into an app, and that works right up until you need a second model, a second tool, an audit trail, or a cost report. At that point the logic is scattered across controllers and cron jobs, and nobody can say with confidence what the system will do with a given request. A clear solution architecture, built around an orchestration layer, is what keeps an AI product maintainable after the demo.

Why AI products need their own solution architecture

A traditional web app is mostly deterministic. The same request hits the same code and returns the same result. An AI feature adds a component that is probabilistic, priced per token, rate limited by a third party, and capable of being talked into things. You cannot treat it like another library call.

That changes what the architecture has to do. It has to route requests to the right model, give the model the right context, control which actions it can take, check what comes out, and record all of it so you can debug and bill. If those concerns live inside individual features, every team reinvents them slightly differently, and the gaps show up as incidents.

The reference architecture below is the one we come back to on most projects, from customer support assistants to internal agents. Not every product needs every box on day one, but you should know where each one would go.

The reference architecture, layer by layer

Think of a request flowing from your app through seven components:

  1. Client apps (web, mobile, WhatsApp, voice) send requests with a user identity and a session.
  2. AI gateway authenticates, rate limits, and routes to model providers.
  3. Orchestration layer runs the workflow or agent loop: builds context, calls models, calls tools, decides when to stop.
  4. Tools are the actions and data sources the model can use, exposed through a controlled interface.
  5. Memory stores conversation state and longer-lived facts about the user or task.
  6. Retrieval finds relevant documents or records to ground the model's answer.
  7. Guardrails and observability wrap everything: checks on input and output, plus traces, metrics and cost data.

AI gateway

The gateway is the single door between your system and model providers. It holds the API keys so application code never does. It enforces per-tenant and per-user rate limits, retries on provider errors, fails over to a backup model when one is down, and records token usage against whoever made the call. It is also where you enforce data residency rules, for example sending requests from Saudi users only to an approved in-region endpoint.

Without a gateway, model keys end up in a dozen services and you cannot answer basic questions like "what did we spend on customer X last month?"

Orchestration layer

This is the core of the system. The orchestration layer takes a request, decides which flow handles it, assembles the prompt and context, calls the model, executes tool calls, and loops or stops. It is where business logic about AI lives: which steps are fixed, which are left to the model, and where a human must approve.

Keep it separate from your product code. Your app should say "handle this support message for this customer" and get back a result and a trace, not build prompts itself. That separation lets you change models, prompts and flows without redeploying every client.

A flow definition might look like this in configuration:

flow: support_reply
model:
  route: classify -> small-fast
  draft: mid-tier
steps:
  - classify_intent
  - retrieve: { index: help_center, top_k: 5 }
  - tools: [lookup_order, check_refund_policy]
  - draft_reply
  - validate: { schema: SupportReply, pii_scan: true }
human_approval:
  when: action in [issue_refund, cancel_order]
limits: { max_steps: 8, max_tokens: 20000, timeout_s: 30 }

Tools

Tools are how the model affects the world: looking up an order, booking a slot, creating a ticket. Each one should have a strict input schema, a narrow purpose and its own permission check that runs with the end user's identity, not a god-mode service account. If you expose tools to multiple assistants, the Model Context Protocol (MCP) is now the common standard and is supported by most major clients. Plan your own audit logging, multi-tenant isolation and rate limiting around it, since the protocol leaves those largely to you.

Memory

Separate short-term memory (the current conversation, trimmed or summarized to fit the context window) from long-term memory (preferences, past cases, facts the user has shared). Long-term memory needs the same care as any user data store: retention rules, deletion on request, and tenant isolation.

Retrieval

Retrieval grounds answers in your own content. A typical setup is a vector index plus keyword search, with results filtered by the user's permissions before they reach the model. That last part is often missed. If a user cannot open a document in your app, the retrieval layer must not quote it to them either.

Guardrails

Guardrails sit on both sides of the model. On the way in: input validation, prompt injection checks on untrusted content, PII redaction. On the way out: schema validation, policy checks, and blocking actions that need approval. They belong in the orchestration layer, applied consistently, rather than sprinkled through features.

Observability

Every request should produce a trace: the prompt, retrieved context, each model call with token counts and latency, each tool call with arguments and results, and the final output. This is how you debug a bad answer, run evaluations on real traffic, and attribute cost per customer or feature. Mask sensitive fields before they hit the logs.

How the orchestration layer keeps things sane

A well-built orchestration layer gives you a few properties that are hard to get any other way:

  • One place to change a model. Swap a provider or tier for one flow without touching app code.
  • Consistent limits. Step caps, token budgets and timeouts apply to every flow, so a confused agent cannot burn through your monthly budget overnight.
  • Replayable runs. Because inputs, context and outputs are recorded, you can replay a failed request against a new prompt and see whether it is fixed.
  • Clear approval points. Human checkpoints are declared in the flow, not hidden in a UI component.
  • Testability. You can run an evaluation set against a flow in CI before it ships.

The most common mistake we see is letting the orchestration layer grow into a second monolith. Keep flows small and composable, and keep domain logic (pricing, eligibility, inventory) in your existing services, exposed as tools. The orchestration layer coordinates. It should not own the business rules.

Build vs buy for each layer

You do not need to build all of this. The right mix depends on how core AI is to your product, your compliance constraints, and your team's capacity.

LayerUsually buy or adoptUsually build
ModelsHosted APIs from major providers; open-weight models if you need self-hostingFine-tunes only when evaluations show a clear gap
GatewayOpen-source or managed AI gatewaysCustom rules for tenancy, residency and billing
OrchestrationA framework or SDK for the agent loopYour flows, approval logic and limits
ToolsVendor MCP servers for common SaaSTools over your own systems and data
RetrievalManaged vector database or search serviceChunking, permission filtering, ranking for your content
ObservabilityLLM tracing and evaluation platformsDashboards tied to your business metrics

A good rule: buy the commodity plumbing, build the parts that encode your product and your risk decisions. Frameworks change quickly, so wrap them behind your own interfaces. If a framework becomes a liability in a year, you want to replace it without rewriting every flow.

Also check where vendors process and store data. For Gulf clients, especially those covered by Saudi PDPL, the hosting region and retention settings of each component matter as much as its features. Our cloud architecture team usually maps this out before any code is written.

A sensible rollout order

You can grow into this architecture in stages:

  1. Start with a gateway and tracing. Even with one feature, centralize keys, usage tracking and logs. It is cheap and saves pain later.
  2. Extract the orchestration layer as soon as you have a second AI feature or a second model. Move prompt building and tool calls out of app code.
  3. Add retrieval and permission filtering when answers need to reference your own content.
  4. Formalize guardrails and evaluations before you give the system write access to anything important.
  5. Add long-term memory last, once you know what is actually worth remembering.

This order means each stage is useful on its own and you are never blocked on a big-bang platform build. It is also the order we typically follow when delivering AI development services for clients who already have a product in production.

How Softzee can help

We design and build AI architecture for products that need to run reliably, from multilingual voice agents to support assistants on WhatsApp, and we can review an existing setup to show where the gaps are. If you are planning your orchestration layer or deciding what to build versus buy, get in touch and we will talk it through.

Solution ArchitectureOrchestration LayerAI ArchitectureLLM GatewayBuild vs Buy

Have a project in mind?

Tell us what you are trying to build. You will get an honest take on scope, timeline and cost, usually within one business day.

Keep reading

All articles