Service — AI Production Audit

Your AI works in a demo. We make it survive production, and cost less doing it.

The AI Production Audit is a fixed-scope engineering audit: you bring the live system or target company, Corteus examines cost, evals, provider exposure, failure points, and data risk, then returns a written report.

Scope the audit Fixed fee, scoped in one call
What it covers

The audit follows the system where the risk lives.

It is technology-agnostic. Orchestration layer, custom pipeline, hosted model, open model, internal gateway: the report reads the architecture the company runs.

01 / Check

Inference cost and model choice

We trace where inference and token cost goes across prompts, tools, retries, context, routing, caching, and model choice. The question: which changes reduce spend while preserving the product behavior that matters.

02 / Check

Eval coverage

We check whether an eval suite exists, whether it represents the real prompts, and how model releases are allowed into production. The question: can the team prove a change holds before users see it.

03 / Check

Provider exposure

We map how tightly the system is bound to one provider across prompts, schemas, hosted features, stored files, and deployment assumptions. The question: what breaks if the preferred provider changes price, policy, or behavior.

04 / Check

Load and human gates

We identify the failure points under real load, queue pressure, retries, tool timeouts, and irreversible actions before a human gate. The question: where does one bad response become an operational event.

05 / Check

Data provenance and license exposure

We inspect source data, derived datasets, retention, embeddings, fine-tuning material, and license terms. The question: which data can be defended when a customer, buyer, or investor asks.

The deliverable

A written report, built to be challenged.

The output is a written report. It uses the same discipline as the build memo: findings stated plainly, counted where numbers exist, specific changes named, and projected effects written as ranges the client can pressure-test.

The report separates evidence from judgment. It shows what to change now, what to measure next, and what remains a business decision.

Read the sample writing standard
The tools run the first pass

Open-source checks create the first evidence set.

The audit starts with wrapper-test, model-rot, and eval-audit. They turn wrapper behavior, model-change exposure, and eval coverage into artifacts a reviewer can inspect.

The audit is the senior read a static tool cannot give: which finding matters, which tradeoff is acceptable, and which system change should come first.

Who it is for

Two readers. One report.

01 / Reader

Companies running AI in production

You have live users, rising inference spend, model-release surprises, eval gaps, or a system that touches decisions with operational consequence.

02 / Reader

Investors evaluating AI startups

You need a pre-investment technical read on unit economics and engineering reality: what the system costs to run, how it is tested, what it depends on, and which risks belong in the term sheet conversation.