Corteus co-builds one AI-native company at a time. This is the thesis we are currently willing to co-build, published so a qualified partner can decide two things before a meeting: whether this company should exist, and whether they are the right person to own it.

The thesis in one sentence. In regulated service operations, the expensive step is not producing an answer. It is proving it.

The costly step is proof, not the answer

Take three markets that look unrelated and are structurally the same: insurance brokerage operations, title and closing work, and clinical or financial administration.

In all three, a professional produces a recommendation. A renewal option. A cleared title. A coverage determination. A suitability note. The recommendation itself is often the short part. What surrounds it is long. Every material recommendation has to carry the sources it rests on, a review by someone qualified to challenge it, an approval by someone accountable for it, and a record that will still make sense to an examiner, an auditor, a plaintiff, or a successor two years later.

That surrounding work is not overhead in the sense of waste. It is the product. A brokerage that gives good advice with no defensible record has not delivered the thing the client is paying for. It has delivered the thing the client would have to defend alone.

Someone pays for this today, and it is usually the most expensive person in the room. The senior operator re-reads the junior’s file. The compliance officer reconstructs why a decision was made. The principal signs and therefore reads. Verification work is not a small support function either. In the United States, claims adjusters, appraisers, examiners, and investigators held 389,700 jobs in 2025 with a median wage of $78,020 in May 2025, and employment is projected to decline about 6 percent through 2035. That is a large population whose output is, essentially, checked judgment, under pressure to shrink.

The market question is therefore not “can software write the recommendation.” Software can write the recommendation. The question is whether software can produce the proof at the same time, to a standard the accountable buyer will sign.

Why generic assistants make this worse

An assistant that returns fluent, confident text with no provenance does not remove verification work. It relocates it, upward, to the person who is least able to do it cheaply.

The reviewer now receives a plausible document with no visible basis. They cannot spot-check it, because there is nothing to check against. They cannot trust it, because they have seen it be wrong. So they redo the underlying work, and then read the draft as well. The pilot reports time saved on drafting. The firm’s actual cost line does not move, and sometimes it worsens.

This is not a claim about weak models. It is what careful measurement of purpose-built tools keeps finding. A preregistered evaluation of leading commercial legal research systems, which use retrieval specifically to ground answers in real authority, found hallucination rates between 17 and 33 percent, published on 30 May 2024. Those are domain products with real document corpora behind them, not chat toys. The federal standards body treats this as a named risk class rather than a defect to be patched: confabulation is one of the risks catalogued in the NIST Generative AI Profile, NIST AI 600-1, released on 26 July 2024.

Read together, those two things say something specific. Grounding helps and does not finish the job. If a system cannot show which source supports which sentence, a regulated firm has to assume any sentence might be unsupported, and the review cost returns in full.

That is the wedge. Not a better writer. A system whose output arrives already carrying the thing the reviewer needs.

Why now

None of the reasons to build this now depend on models getting better. They are all about what buyers are being required to be able to show.

The securities regulator has already written the expectation down. The 2026 FINRA Annual Regulatory Oversight Report, published in December 2025, describes generative AI practices at member firms in terms of storing prompt and output logs, tracking which model version was used and when, and human review of model outputs. Those are not features. They are records a firm must be able to produce.

Insurance supervision has moved the same way, state by state. The NAIC model bulletin on the use of AI systems by insurers was adopted on 4 December 2023, and the association’s own implementation map, status as of 6 August 2026, lists 25 adopting jurisdictions plus separate insurance-specific rules in California, Colorado, New York, and Texas. Documented governance of AI-assisted decisions is now the default expectation for a large share of the US market, not an early-adopter posture.

Health care administration has hard dates attached. Under the CMS interoperability and prior authorization final rule, payers must give a specific reason for denied prior authorization decisions beginning in 2026, post prior authorization metrics publicly, and meet the API requirements by 1 January 2027. A decision that must carry a specific stated reason is a decision that needs traceable evidence behind it. Separately, certified health IT that supplies predictive decision support must publish nine categories of source attributes covering purpose, out-of-scope use, development data, fairness process, external validation, performance, and ongoing maintenance. That is a public schema for what “we can explain this system” means.

Europe is worth reading carefully, because the honest version of the story is less dramatic than the marketing version. The AI Act requires automatic logging and record-keeping for high-risk systems under Article 12, and requires deployers to assign oversight to people with the competence and authority to exercise it, and to keep logs for at least six months, under Article 26. But the deadline moved. The AI Omnibus entered into force on 27 July 2026 and pushed application of the Annex III high-risk rules to 2 December 2027. A partner should not be sold urgency that is not there. The requirement is coming, the panic is not, and a company that only sells deadline fear will be exposed when the deadline slides again.

Meanwhile adoption is running ahead of the evidence layer. In US Census Bureau survey data, AI use in finance and insurance stood at 33.9 percent against a national rate of 19.8 percent as of 3 May 2026. Regulated firms are adopting faster than average while the obligation to explain what the systems did is being written around them. That gap is the opportunity, and it is also the risk: it closes either by firms retreating from AI, or by someone building the layer that makes the use defensible.

What evidence must exist before anyone builds

Corteus does not start with a build. It starts with a bounded validation that can fail. For this thesis, five artifacts have to exist before product investment is justified, and four of them come from the partner’s own systems rather than from anyone’s opinion.

Countable volume. How many times per month does this exact decision happen, pulled from the firm’s records rather than recalled from memory? A workflow that happens irregularly cannot amortise an evidence pipeline.

A named cost. Whose hours, at what seniority, currently go into re-checking? “It takes a lot of time” is not a number. “Two senior reviewers, roughly a day a week each, on this queue” is.

Real source material. A sample of actual files, in the state they actually arrive in, including the scanned fax, the inconsistent carrier form, and the document someone rotated by mistake. Sample quality decides feasibility more than model choice does, and it is the cheapest thing to test first.

A written definition of acceptable evidence. From the person who signs, not from the person who is enthusiastic. What must a recommendation show for them to approve it without redoing the work? If they cannot write that down, the product has no acceptance criteria and the project has no end.

A baseline. Current cycle time, current rework rate, current audit exception rate. Without a before, the after is a story. This is the same discipline as a real eval suite, applied to the business rather than the model, and it is what separates an agent that is actually working from one that is merely running.

What the domain partner brings

The qualification is the same one on our venture page, stated plainly.

Firsthand domain knowledge. Not interest in the sector. Time inside it. Knowing which exception eats the afternoon and why the obvious fix has already failed twice.

One real, costly, testable problem. Specific enough to instrument. If it cannot be measured, it cannot be validated, and a build against it is a guess with a budget.

Customer access. A path to the buyers who own this workflow, through a book of business, an operating role, a distribution relationship, or a design-partner firm that will actually put files in front of us.

Committed capital or a credible funding path. Corteus is not a fund and does not supply capital. Validation and a first production build have to be financed.

A full-time operating owner. One person who owns the market and the company, not a sponsor who checks in. This is the qualification that most often decides whether a conversation becomes a company.

An idea, on its own, is not any of these.

What Corteus commits

Validation. Framing the falsifiable problem, running the bounded test, and writing the decision memo including the case for stopping. Stopping early is a successful outcome of validation, not a failure of it.

Product and production-AI engineering. The smallest production system that proves the operating thesis, built so the evidence, the permissions, the audit record, and the human approval are part of the system rather than a policy document beside it.

The initial technical founding capability. The technical founding team for the first phase, so the operating owner does not have to hire an engineering organisation before knowing whether the company should exist.

The relationship is a venture partnership. Corteus is not a fund, an agency, or a bench of contractors, and it does not build for free against an unvalidated idea.

How we work, which is not the same as a result

Corteus is early. We have no portfolio outcome, no client result, and no revenue figure to put here, and a memo that implied otherwise would fail its own thesis.

What we can show is method, because our own operating system is built the way we would build this company. Model judgment is separated from a deterministic control plane that owns permissions, state, budgets, and approvals. Actions are recorded in an immutable audit log. Any action that leaves the system requires exact human approval first. Work stops when a claim cannot be supported, and the stop is itself recorded.

That last one is worth stating precisely, because it is the difference between the discipline being real and being a slogan. Our own outbound work has been halted by that rule, on an unsupported claim, and the halt is in the record rather than in a sentence like this one. The artifacts are public.

Frame this correctly. It is how we operate. It is not evidence that a customer got a result, and it should not be read as one.

A hypothetical, clearly labeled

The following is hypothetical. It is not a client, a project, a pilot, or a result. It is included to make the shape of the company concrete.

Suppose a commercial insurance brokerage handles a recurring renewal review. Today an account manager assembles the file, drafts the recommendation, and a principal reviews it, mostly by re-reading the underlying policy wording because the draft asserts coverage positions without showing where they come from.

The evidence-native version does not write a better recommendation. It produces the recommendation with each coverage assertion bound to the clause it came from, flags the assertions it could not bind, routes those to a person by name, records the approval, and retains the whole chain in a form an examiner could read. The principal’s job changes from reconstruction to adjudication. What would be measured is not words generated. It is reviewer minutes per file, the share of assertions that arrive already bound to a source, and the audit exception rate.

Whether a brokerage would pay for that, and how much, is exactly what validation is for. We do not know it yet, and neither does anyone else selling into that room.

What we do not know

Four unknowns are genuinely open, and a partner should press on them rather than accept them.

Whether the accountable buyer can actually write down what evidence they would accept. Many cannot, on first asking. If that never resolves into something testable, there is no acceptance criterion and no product.

Whether the source material survives contact. Legacy document estates in these markets are worse than they look in a demo. This is a feasibility question, not a modelling question, and it is answered with real files or not at all.

Whether the evidence layer is a company or a feature. The incumbent system of record may absorb it. A defensible answer needs a wedge that the incumbent cannot copy cheaply, usually proprietary workflow data, a distribution position, or an accountability relationship rather than a clever interface.

Whether firms will pay for provability separately from headcount reduction. Regulated buyers say they value defensibility. Budgets sometimes only respond to labour savings. If the only purchase justification is cutting heads, the economics are different and the sales motion is different.

When we stop

The stop conditions are written before the build, and they are not negotiable afterwards.

We stop when there is no repeated volume, because a bespoke decision cannot support the cost of a proof pipeline. We stop when nobody can name a measurable cost, because a saving that cannot be measured cannot be sold twice. We stop when the source material is unusable, because a citation to an unreadable document is not evidence. We stop when there is no accountable buyer, because a system built to satisfy scrutiny needs someone who is actually subject to it.

Any one of these ends the work. That is the point of writing them down in public, before a partner conversation, rather than discovering them at month five.

If you own one of these workflows

This memo is a call for one partner, not a service offering. If you operate inside a regulated services firm, own or can reach a repeated evidence-heavy workflow, and are prepared to own the resulting company full time, that is the conversation.

Bring the workflow, one month of real volume, and the name of the person who signs. We will bring the validation and, if the evidence holds, the technical founding capability.

Co-build with Corteus

Sources: bls.gov, Claims Adjusters, Appraisers, Examiners, and Investigators, modified 27 August 2026 | arxiv.org, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, 30 May 2024 | nist.gov, NIST AI 600-1 Generative AI Profile, 26 July 2024 | finra.org, 2026 Annual Regulatory Oversight Report, GenAI, December 2025 | naic.org, Implementation of NAIC Model Bulletin, status as of 6 August 2026 | cms.gov, Interoperability and Prior Authorization Final Rule CMS-0057-F fact sheet, 17 January 2024 | healthit.gov, 170.315(b)(11) Decision Support Interventions test method | europa.eu, EU AI Act Article 12 record-keeping | europa.eu, EU AI Act Article 26 deployer obligations | europa.eu, AI Omnibus enters into force, 27 July 2026 | census.gov, Large Firms With at Least 20 Employees Biggest AI Users, 26 May 2026