AI Benchmarking / Cortega RedTeam

Download the benchmark.
Run it whenever you need.

Cortega turns public datasets and current AI research into benchmarking software you run yourself. Install it on-premises, choose your models and LLM judges, and test every change—without commissioning a new assessment each time.

Downloadable productRegular release updatesAutomated LLM judging
Your own benchmark capability

A benchmark should be part
of how you build AI.

Models change. Prompts change. Guardrails change. Cortega gives your team a way to test those changes as often as the work demands, using software installed in your environment.

RUN IT YOURSELF

Test on your schedule.

Run a baseline today, check a new configuration tomorrow, and repeat before release. You control when the evaluation happens.

KEEP IT CURRENT

Get research as runnable software.

Regular Cortega releases bring updated benchmark implementations. Your team can use new methods without building the test harness from scratch.

PREPARE WITH EVIDENCE

Find weaknesses before an external assessment.

Use internal runs to identify failures, improve the system, and retest before bringing in an independent evaluator.

From research to repeatable execution

We build the benchmark software.
You run the benchmarks.

Public datasets and the latest scientific research in LLM evaluation and security become executable suites in Cortega. Cortega RedTeam brings the test harness, attack workflows, scoring, and run history together.

  1. Download and install

    Deploy Cortega on-premises and receive updates through regular product releases. Benchmark execution is available only in the on-premises solution.

  2. Choose the test and judges

    Select a suite, the model or protected configuration to test, and the LLMs used to judge responses. Use model access you control and pay for.

  3. Run, inspect, repeat

    The software executes the tests and scores responses automatically. Review the evidence, change your configuration, and rerun the same suite.

No human grading step in the run.

Cortega’s execution and LLM judging are automated. You do not send each run to a panel of reviewers or wait for a service provider to grade it. Your team chooses the models, pays for their usage, and decides what to do with the results.

Repeatable execution means the same suite and recorded setup can be rerun. Model responses and LLM judgments can vary; inspect the judgments and compare repeated runs when making a release decision.

What you can test

One product.
More than a set of prompts.

Cortega packages the execution needed for different tests: fixed tasks, evolving attacks, tool interactions, and comparisons with protections enabled. Each suite answers a specific question about the system you are building.

LLM & DOMAIN BENCHMARKS

Can the model do the task correctly?

Evaluate contract provisions, claims against contract text, legal citations, and domain-specific safety scenarios. Use the same tasks and scoring method to compare candidate models.

Useful before choosing a model or changing a prompt.

ADAPTIVE SECURITY TESTS

Does the defense survive a persistent attacker?

Run single-shot, many-shot, and multi-turn attacks. Dynamic tests use an attacker LLM that reacts to the target’s replies and changes its next attempt.

Useful when a single refusal does not tell you whether a defense holds.

MCP & TOOL TESTS

Can a tool result steer the agent off course?

Test poisoned MCP results and injected tool-call instructions. Evaluate behavior beyond a direct chat prompt, where agents encounter content from other systems.

Useful before connecting agents to tools and external content.

GUARDRAIL & BOOSTER COMPARISONS

What does the protection actually add?

Compare the model alone with guardrails or verification enabled. Check attack outcomes, harmless-request refusals, citation quality, cost, and latency according to the suite.

Useful before approving a security policy or adding a booster.

Cortega’s advantage is the complete workflow: research implemented as software, automated execution and judging, and evidence you can revisit after each change.

Explore public results, datasets, and methods →

See the benchmark architecture and execution modes ↓

Inside the benchmark runner

Cortega RedTeam architecture

Run industry, security, and MCP benchmarks against a selected gateway and model. Compare observed behavior with expected outcomes and retain the evidence for each run.

Industry, security and MCP suites feed the benchmark runner. Test traffic goes through a selected gateway and model, and results return for scoring. Dynamic multi-turn tests use response judgments to generate follow-up attacks. Scores, sample evidence and configuration are saved in run history.
Teal: test requests and follow-up turns. Gray: returned results. Dashed: saved results. Benchmark traffic is separate from production traffic; runs do not change guardrail configuration.
Execution modes and scoring

Request-guardrail mode stops before the model. LLM-then-drop mode includes the model but stops before response checks. Full flow includes request checks, the model, and response checks. Mode availability depends on the suite; adaptive tests need model responses.

Attack and benign samples measure both missed threats and over-blocking. The legal citation suite checks returned citations. MCP probes test poisoned tool results; full flow also checks whether the model proposed the injected tool call.

Security is also an availability question

One compromised agent.
How many applications lose access?

If an attacker steers your chatbot or agent into sending abusive requests, that activity reaches the model provider under your account. Providers can restrict or suspend access for policy violations or security incidents. If many agents share that dependency, the disruption can reach beyond the attacked application.

TEST THE BOUNDARY

Find what gets through before an attacker does.

Use adaptive attacks and MCP tests to discover whether malicious inputs can bypass your configured defenses. Test request guardrails to see what they stop before content reaches a hosted model.

PLAN FOR CONTINUITY

Understand the dependency your agents share.

Review which applications rely on the same provider account. Use benchmark evidence to improve protections and assess approved alternatives, including privately hosted models, before a disruption occurs.

A malicious request does not automatically mean account suspension, and benchmarking cannot guarantee uninterrupted access. On-premises execution also does not remove provider-policy exposure if you use hosted targets or judges.

Provider policy context

Provider policies allow account enforcement for misuse or security events. See OpenAI’s account deactivation guidance and the OpenAI Services Agreement. The wider impact on applications sharing an account is an operational dependency to assess.

Make the next run useful

Find the failure.
Fix it. Run it again.

An overall score tells you where to look. Individual responses, judgments, and category results help you understand what failed. Run history lets you compare the next version with the baseline.

RedTeam · Benchmark score
RedTeam: overall score plus a pass rate for each category, with precision, recall, and F1
RedTeam: every benchmark prompt with its category, expected and actual outcome, and latency

From the Cortega console · red-team and finance runs.

Our strategy: make each new test easier to run

Invest in the capability.
Get more value from every run.

An external assessment gives you an outside view at a point in time. Cortega gives your team reusable benchmarking software for all the changes in between. Run internal checks often, fix the weaknesses you find, and bring in independent evaluators when you need their judgment.

REUSE THE SOFTWARE

Less effort to start the next test.

Use the installed product to evaluate another model, prompt, or policy. Routine reruns do not require commissioning a new service engagement.

AUTOMATE THE RUN

Spend less time executing and grading.

The harness runs the suite and your selected LLM judges score the responses. Your team can focus its attention on failures and improvements rather than manually processing every answer.

GET RESEARCH THROUGH RELEASES

Avoid rebuilding the benchmark stack.

Cortega turns public datasets and research methods into maintained, runnable suites. Product updates reduce the work of assembling and maintaining those test workflows yourself.

PREPARE BEFORE YOU COMMISSION

Make external assessment time count.

Find and fix weaknesses internally before seeking independent review. Use specialist assessment for outside judgment, certification, and the questions that need human expertise.

You control the cost of running the tests.

Choose the target models and LLM judges, decide how often to run, and compare cost alongside results. Budget for the Cortega product, your infrastructure, and model usage. Software reuse and automated grading can lower the total effort and cost of frequent benchmarking; the economics depend on your run volume and model choices.

Cortega complements independent assessment. Internal scores do not confer certification or predict an external score, and LLM judgments do not replace professional review where it is needed.

MAKE BENCHMARKING A TEAM CAPABILITY

Your environment.
Your judges. Your schedule.

Download Cortega. Build a baseline. Be ready for the next change.