The Fight Between Attackers and Your Chatbot Is Completely Asymmetric — Cortega Blog
Blog · Part 1 of 2

The Fight Between Attackers and Your Chatbot Is Completely Asymmetric.

Sridhar Ramachandran · CEO & Co-founder · September 2026

There's been non-stop talk about AI guardrails over the last few days. Most of it is about the threat of AI going rogue and jumping its own guardrails. The flip side gets less airtime, but it's just as real: why do models keep getting tricked? Why does prompt injection feel so hard to solve? Why do standard, filter-like guardrails collapse the moment they hit production traffic?

I've always maintained that the nature of AI's advancement makes this fundamentally asymmetric. Here's the economic and security reality behind one of the most common enterprise use cases there is: the AI-based customer service chatbot.

Guardrails coverage is asymmetric: a wide bar for policy and alignment statements next to a short bar for runtime enforcement, every request, every session, with the Cortega mark on the short bar.

The fight between attackers and enterprise chatbots is completely asymmetric. Here's the dynamic:

1. The enterprise reality: unit economics

When companies run AI-based customer support or automated workflows at scale, they typically reach for fast, economical models — a Claude Haiku, a GPT-4o-mini, a Gemini Flash — priced around $0.001 or less per query. That's the right call: it doesn't make economic sense to route every routine support ticket through a $30-per-million-token heavy reasoning model.

2. The attacker's advantage: compute disparity

Attackers don't operate under that same unit-economic constraint. An adversary will happily spend $0.50 running an autonomous bot powered by a heavyweight frontier reasoning model directly at a company's customer support bot — pocket change, if it extracts a $400 refund or manipulates an internal tool on the other end.

That's a real threat, and it's structural, not incidental: the enterprise side is optimizing for cost per query at scale, and the attacker side is optimizing for success per attempt. Those are different objective functions, and the second one can afford to spend far more per interaction than the first one budgeted for.

So what should organizations do? The guardrails doing the defending can't be built on the same cost assumptions as the system they're defending. They need to sit in-path, external to the model being attacked, and hold the line regardless of which model is generating the attack on the other side. Easier said than done.

Part 2 is here: three exact attack vectors frontier bots use against customer-facing agents — from fake manager overrides to auth-payout traps — run live in the Cortega Arena, standard guardrails versus Cortega Policy Guardrails.

What's your team's approach to securing customer-facing AI agents today?

Originally posted on LinkedIn. · Back to all posts

Get started

Break the asymmetry

See it hold the line

Cortega's AI Border Gateway inspects and enforces policy on every request, in-path, before it reaches a model.

Test your guardrails

Run a suite of unsafe prompts — including fraud enablement — through the guardrails you have live right now.