How to Implement Guardrails for AI Apps (and Why Hack Proof Is a Myth)

How to implement guardrails for LLM apps: prompt injection defenses, tool permissions, schema validation, red teaming, and why no AI system is truly hack proof.

S
Softzee EngineeringSeptember 19, 2026 · 7 min read

Every AI feature that reads untrusted text, whether a customer message, a web page, an uploaded PDF or an email, can be instructed by that text. If the same feature can also call tools or see private data, an attacker does not need to break your servers. They just need to write a convincing sentence. This article covers how to implement guardrails that make that attack hard and limit the damage when it works, and why "hack proof code" is the wrong goal.

Start with an honest threat model

Let us be direct: there is no such thing as a hack proof LLM application. Language models do not separate instructions from data the way a SQL engine separates a query from its parameters. Everything in the context window is just text the model weighs. Vendors keep improving resistance, and you should use those improvements, but no current model or filter blocks prompt injection reliably on its own.

So the useful question is not "how do we make it impossible?" It is "what can this system do if the model is fully convinced to misbehave, and how do we make that outcome small, visible and recoverable?" That reframing drives every decision below.

For each AI feature, write down three things:

  • Untrusted inputs. Anything the model reads that a third party could have written: user messages, retrieved documents, tool results, emails, file uploads, web content.
  • Sensitive assets. Data the model can see (customer records, internal docs, system prompts) and secrets it should never reveal.
  • Capabilities. Every tool or action it can trigger, especially ones that send data out (email, webhooks, links, images) or change state (refunds, bookings, deletes).

The dangerous combination is all three at once: untrusted input, access to private data, and a way to send that data somewhere. If a feature has all three, it needs the strongest controls, or it needs to be redesigned so it does not.

Prompt injection: direct and indirect

Direct injection is the user typing "ignore your previous instructions" or a more creative variant. It is the version everyone tests, and it is the less dangerous one, because the attacker is usually only attacking their own session.

Indirect injection is worse. The malicious instruction sits inside content the model reads on someone else's behalf: a support ticket that tells the agent to forward the customer list, a web page that tells a research agent to visit a URL with data in the query string, or a CV with hidden text telling the screening assistant to rate the candidate highly. The legitimate user never sees it.

Practical mitigations that help, none of which are sufficient alone:

  • Clearly mark untrusted content in the prompt (for example, wrap it in tagged blocks) and tell the model to treat it as data only.
  • Run a classifier on untrusted content to flag likely injection attempts before the main model sees it.
  • Keep the system prompt free of secrets. Assume it will leak eventually.
  • Strip or neutralize channels that can exfiltrate data, such as auto-loading images or links built from model output.

Input and output filtering

On the way in

Input filters reduce noise and catch obvious attacks. Validate length and format, reject unexpected file types, detect and redact PII you do not need, and screen for known injection patterns. Treat these as speed bumps. An attacker who wants to get past a keyword filter will, so do not let your safety depend on it.

On the way out

Output filtering is often more valuable, because it checks what the model is actually trying to do. Scan responses for sensitive data that should not leave (account numbers, internal URLs, other customers' details), for policy violations, and for content that does not match the task. For customer-facing chat, a lightweight second model pass that checks "does this reply only discuss this customer's own order?" catches a surprising number of problems.

Tool permissioning is your strongest control

If you only invest in one guardrail, make it this one. The model's permissions should never exceed the permissions of the person it is acting for, and ideally should be narrower.

  • Run tools as the user. Pass the end user's identity through to every tool call and enforce authorization in the tool, not in the prompt. "Only look up the current customer's orders" written in a prompt is a suggestion. The same rule enforced in the API is a control.
  • Read and write separation. Give the model read-only tools by default. Write actions get separate tools with extra checks.
  • Human approval for high-impact actions. Refunds, cancellations, outbound messages to third parties and anything irreversible should require a confirmation outside the model's control.
  • Allowlists for outbound destinations. If an agent can send email or call URLs, restrict where. An injection that cannot reach an attacker-controlled address is much less useful.
  • Scoped, short-lived credentials. Avoid static API keys with broad access. Issue tokens per session with only the scopes that flow needs.

This also applies to tools exposed over MCP. The protocol standardizes how tools are described and called, and remote servers use OAuth 2.1, but static secrets are still common in practice and audit trails and per-tenant isolation are mostly left to you. Review every MCP server you connect as you would any third-party integration.

Validate outputs with schemas, not hope

When the model's output drives code, such as a tool call, a database write or a UI update, never parse free text and trust it. Ask for structured output and validate it strictly. Reject anything that does not match, and treat validation failures as signals worth logging.

from pydantic import BaseModel, Field, ValidationError
from typing import Literal

class RefundDecision(BaseModel):
    order_id: str = Field(pattern=r"^ord_[a-z0-9]{10}$")
    action: Literal["approve", "deny", "escalate"]
    amount: float = Field(ge=0, le=500)   # hard cap, larger goes to a human
    reason: str = Field(max_length=300)

try:
    decision = RefundDecision.model_validate_json(model_output)
except ValidationError:
    decision = escalate_to_human(model_output)

# Business rules still run in code, not in the prompt
if decision.order_id not in orders_for(current_user):
    raise PermissionError("order does not belong to this user")

Notice the last check. Schema validation confirms the shape. Your own code still confirms the output makes sense for this user and this request. The model proposes; deterministic code decides.

Red teaming and continuous testing

Guardrails rot. A prompt change, a model upgrade or a new tool can quietly reopen a hole you closed months ago. Treat adversarial testing as part of your test suite, not a one-time audit.

  1. Build an attack set. Collect injection attempts, jailbreaks, data extraction prompts and abuse cases relevant to your product, in every language you support. For Gulf products that means Arabic and English at minimum, plus mixed-language inputs, since filters tuned on English often miss Arabic phrasing.
  2. Run it in CI. Every prompt or model change runs against the attack set, and regressions block the release.
  3. Do periodic manual red teaming. Automated sets catch known patterns. People find new ones. Give internal testers or an external team time and a clear scope to break the system.
  4. Feed production back in. When monitoring flags a real attempt, add a sanitized version to the attack set.

Our QA automation and testing team builds these adversarial suites alongside normal functional tests, so security checks run on every release.

Defense in depth, put together

No single layer here is reliable. Stacked together, they make an attack much harder to pull off and much easier to spot.

LayerWhat it stopsWhat it misses
Input filteringObvious injections, unneeded PII, malformed inputNovel or obfuscated attacks
Prompt structureSome confusion between data and instructionsDetermined injection
Tool permissioningAccess beyond the user's rights, unsafe destinationsMisuse within the user's own rights
Human approvalHigh-impact actions going through uncheckedReviewer fatigue if overused
Output validationMalformed or out-of-policy actions and data leaksSubtle but valid-looking errors
Monitoring and alertsAttacks in progress, unusual tool patternsNothing, if nobody reads them

Log every model call and tool call with the user, inputs, outputs and decision, with sensitive fields masked. Alert on patterns like a spike in validation failures, repeated denied tool calls, or outbound requests to new domains. When something does get through, and eventually something will, those logs are how you scope the incident in hours instead of weeks. If you are designing these controls into a new agent, our AI agent development work treats them as part of the architecture rather than an add-on.

How Softzee can help

We build and review AI systems with these controls in place, from WhatsApp support assistants to agents that act on business systems, and we can red team an existing feature to show where it is exposed. If you want a second pair of eyes on your guardrails, reach out and we will start with your threat model.

Implement GuardrailsHack Proof CodePrompt InjectionAI SecurityRed Teaming

Have a project in mind?

Tell us what you are trying to build. You will get an honest take on scope, timeline and cost, usually within one business day.

Keep reading

All articles