Solutions / Private AI Infrastructure

Private AI Infrastructure

Choose where your firm’s AI models run: a local machine, your own GPU server, your cloud account, or dedicated rented GPUs. Cortega connects legal AI tools to those models and applies access, routing, and data policies.

Why private AI

The problem

When AI runs on a hosted service, every prompt, document, and answer is processed on infrastructure someone else operates. Private AI runs the models where you control the data, the access, and the cost.

Today

Prompts go to a hosted AI provider

  1. Your tools and agentsChat, code, documents
  2. Source codeCustomer recordsContractsCredentials
  3. Hosted AI providerOutside your infrastructure
  4. Data is processed where you do not control it
  5. Provider terms set location and retention
  6. The provider chooses and updates the model
  7. Per-token cost grows with every user
With private AI

Prompts stay on infrastructure you control

  1. Your tools and agentsSame tools, one gateway
  2. AI Border GatewayIdentify · Govern · Route
  3. Your model hostLocal, GPU server, or your cloud
    Your recordsLogs and stores you hold
  4. Inference and gateway data stay in your boundary
  5. You set access, routing, and data policy
  6. You choose the model and keep it
  7. Cost follows your own capacity

What you get

Your infrastructure
ToolsGatewayModel

Keep data inside

ProblemEvery prompt and document sent to a hosted service is processed on someone else’s infrastructure.

CortegaRun inference and the gateway on infrastructure you control.

A. Local machineYours
B. Firm GPU serverYours
C. Your cloud accountYours
D. Rented GPUsProvider

Choose where it runs

ProblemRegulation, contracts, and client commitments can restrict where data may be processed.

CortegaPick a local machine, your GPU server, or your cloud account. Rented GPUs are available when provider hosting is approved.

Compare the four options ↓
Sensitive workPrivate model
Standard workApproved model

Route by sensitivity

ProblemNot every request needs the same model, and sensitive work should never default to a hosted one.

CortegaSend sensitive work to private models and standard work to approved models.

Legal teamPrivate model
EngineeringPrivate + approved
ContractorsNo access

Control access

ProblemA shared model needs to know who is using it and how much.

CortegaSet model access, usage records, and spending limits per team.

Open-weight modelLicense checked
OllamavLLM
Chosen with Model Intelligence

Choose the model

ProblemA hosted provider decides which model you get and when it changes.

CortegaSelect an open-weight model for the work and run it until you decide to change.

Hosted API: grows with useYour GPUs: fixed capacity

Control costs

ProblemPer-token pricing grows with every user and every document.

CortegaRun open-weight models on your own GPU capacity. Costs depend on utilization.

1Local machinePilot
2Firm GPU serverShared
3Your cloud accountScale

Start small and grow

ProblemCommitting to hardware before you know the workload is a risk.

CortegaStart with one machine, then move to a shared GPU server or your cloud account. The gateway and policies stay the same.

Your GPU serverDaily use
Rented GPUsOne batch · data leaves

Add capacity for batches

ProblemA large document batch can need more GPUs than you run day to day.

CortegaRent dedicated GPUs for the job where provider hosting is approved. You pay for hours while they run.

Primary: your vLLM serverApproved
Fallback: your cloud GPUsApproved
Fallback: hosted APINot approved

Approve every fallback

ProblemA fallback to a hosted model can send data somewhere you did not approve.

CortegaName the primary and fallback destinations explicitly. Every fallback follows the same data policy.

Private here means inference and gateway data stay within your chosen infrastructure. Agent telemetry, external tools, web search, and hosted fallback models need separate review. Option D runs on a provider’s infrastructure, so data leaves yours.

Deployment options

Four ways to run your models

The main decisions are where documents are processed, who operates the hardware, and whether you need continuous capacity or a temporary batch. The examples below describe possible fits; model quality and capacity depend on the workload.

OptionWhere it runs & howUseful for a legal teamCost & responsibility
A. Local machineA local open-weight model served by Ollama on a machine you control.A single-lawyer pilot; testing drafting, summarization, or document questions with a small model.Uses available hardware. Memory limits model size and context; you maintain the machine and runtime.
B. Firm-owned GPU servervLLM on an existing compatible GPU host, reached through the firm’s network.Shared inference for sustained document review and daily legal AI use.Hardware purchase or existing capacity, power, and operations. The firm manages capacity, security, and updates.
C. GPUs in your cloud accountvLLM in a cloud environment you administer. Cortega’s provisioning scripts support AWS; other clouds use an existing compatible host.Shared firm access without buying a server, with network, region, and storage choices in your account.Cloud compute, storage, and networking charges. You manage the account, access policies, and running resources.
D. Dedicated rented GPUsYour selected model on dedicated capacity operated by a GPU rental provider.Temporary capacity for a large discovery or document-review batch, where provider-hosted processing is approved.GPU-hour billing while running, including idle time. Data leaves your infrastructure; provider location, retention, and terms matter.

Options A–C keep inference in infrastructure you control when the gateway, stores, and auxiliary services are configured there too. Option D is dedicated capacity on a provider’s infrastructure. Dedicated hardware alone does not make it equivalent to an on-premises deployment.

Architecture

Tools, gateway, and model hosts

Your tools call AIBG. The gateway identifies the caller, applies configured controls, and routes to an approved model endpoint. A deployment can use one option or several; every fallback must follow the same data policy as the primary route.

Legal tools call AI Border Gateway inside the firm's chosen infrastructure. It routes to Ollama on a local machine, vLLM on a firm GPU server, or vLLM in the firm's cloud account. An optional route crosses that boundary to provider-hosted dedicated rented GPUs. Gateway records stay in configured governance stores.
Teal: governed requests. Dashed amber: optional provider-hosted processing outside the firm’s infrastructure. Responses return through the gateway.

Keep inspection models, logs, backups, and verification services inside the intended boundary too. Web search, MCP tools, extensions, and external fallback models can send data elsewhere even when the main model is private.

Setup

From a model host to a legal workflow

01

Choose a model and host

Match model quality, document length, memory, and concurrent users to the work. Start Ollama locally, deploy vLLM to an existing server or your AWS account, or create a dedicated deployment with a rental provider.

02

Register the endpoint in AIBG

Add the provider and model, store its credentials in the gateway, and configure team access. Place AIBG and its data stores in the intended environment. Approve primary and fallback destinations explicitly.

03

Connect the tools

Use Cortega Legal for the legal VS Code workspace, Endpoint Guard to configure supported assistants, or client-specific gateway settings. Endpoint Guard’s separate app-inspection path still sends traffic to its original destination; it does not replace a SaaS model with a private one.

04

Test the actual workflow

Run representative documents, streaming requests, and tool calls. Check the selected model, failure routes, and captured data. Use Red Team to evaluate the chosen guardrails and model before rollout.

Red Team →
Step-by-step deployment guide →
Packages & operations

Separate model hosting from gateway management

Foundation

Connect self-hosted models and registered rented GPU endpoints to AIBG. Model routing, guardrails, MCP governance, observability, and cost tracking are included. Hardware and provider charges are separate.

Enterprise

Add rented GPU lifecycle controls: attach an existing deployment, start or stop it, stop idle capacity, and track estimated spend and limits. Creating, sizing, and deleting the deployment remains with the provider. Enterprise also adds platform access and operational controls.

Cortega Hosted is a separate hosted service offering; it is not a model installation in your own cloud account. Air-gapped operation is a separate deployment choice that requires local runtime dependencies, models, and services.

Compare packages →

Private AI Infrastructure

Plan your firm’s deployment

Start with the documents, users, data-location requirements, and expected workload.