Your AI works in a demo. We make it survive production, and cost less doing it.
The AI Production Audit is a fixed-scope engineering audit: you bring the live system or target company, Corteus examines cost, evals, provider exposure, failure points, and data risk, then returns a written report.
The audit follows the system where the risk lives.
It is technology-agnostic. Orchestration layer, custom pipeline, hosted model, open model, internal gateway: the report reads the architecture the company runs.
Inference cost and model choice
We trace where inference and token cost goes across prompts, tools, retries, context, routing, caching, and model choice. The question: which changes reduce spend while preserving the product behavior that matters.
Eval coverage
We check whether an eval suite exists, whether it represents the real prompts, and how model releases are allowed into production. The question: can the team prove a change holds before users see it.
Provider exposure
We map how tightly the system is bound to one provider across prompts, schemas, hosted features, stored files, and deployment assumptions. The question: what breaks if the preferred provider changes price, policy, or behavior.
Load and human gates
We identify the failure points under real load, queue pressure, retries, tool timeouts, and irreversible actions before a human gate. The question: where does one bad response become an operational event.
Data provenance and license exposure
We inspect source data, derived datasets, retention, embeddings, fine-tuning material, and license terms. The question: which data can be defended when a customer, buyer, or investor asks.
A written report, built to be challenged.
The output is a written report. It uses the same discipline as the build memo: findings stated plainly, counted where numbers exist, specific changes named, and projected effects written as ranges the client can pressure-test.
The report separates evidence from judgment. It shows what to change now, what to measure next, and what remains a business decision.
Read the sample writing standardOpen-source checks create the first evidence set.
The audit starts with wrapper-test, model-rot, and eval-audit. They turn wrapper behavior, model-change exposure, and eval coverage into artifacts a reviewer can inspect.
The audit is the senior read a static tool cannot give: which finding matters, which tradeoff is acceptable, and which system change should come first.
Two readers. One report.
Companies running AI in production
You have live users, rising inference spend, model-release surprises, eval gaps, or a system that touches decisions with operational consequence.
Investors evaluating AI startups
You need a pre-investment technical read on unit economics and engineering reality: what the system costs to run, how it is tested, what it depends on, and which risks belong in the term sheet conversation.