Softzee blog

Notes on building AI that works in production

Practical writing from our engineers on agentic AI, RAG, MCP, AI security, compliance and the cloud infrastructure underneath it all.

AI Engineering
Oct 6, 2026 · 7 min read

Save Token Cost: 7 Practical Ways to Cut Your LLM Bill in Production

How to save token cost on LLM apps: prompt caching, model routing, context trimming, batching, output limits, answer caching and measuring cost per task.

Read article
AI Engineering
Oct 3, 2026 · 8 min read

Agentic Engineering: Picking the Right Agent for the Right Job

A practical guide to agentic engineering: when to use one agent or many, how to pick models per task, design tools, and place human checkpoints.

Read article
AI Engineering
Sep 29, 2026 · 6 min read

RAG in Production: Retrieval Augmented Generation Done Properly

RAG that works for real users: chunking, hybrid search, reranking, permission-aware retrieval, freshness, evaluation, and when to skip RAG entirely.

Read article
AI Engineering
Sep 26, 2026 · 7 min read

Solution Architecture for AI Products: The Orchestration Layer

A reference solution architecture for AI products: gateway, orchestration layer, tools, memory, retrieval, guardrails and observability, plus what to build or buy.

Read article
AI Engineering
Sep 22, 2026 · 6 min read

MCP Explained: A Practical Guide to the Model Context Protocol

What MCP is, how it differs from function calling and APIs, how to build MCP servers for internal systems, and how to secure them with OAuth and audit logs.

Read article
AI Engineering
Sep 10, 2026 · 6 min read

Frontline AI Agents: Putting AI in Front of Your Customers

What it takes to put AI on the frontline with customers: voice agents, WhatsApp and web assistants, Arabic and English, human handoff, SLAs and what to automate first.

Read article

Have a project in mind?

Tell us what you are trying to build. You will get an honest take on scope, timeline and cost, usually within one business day.