An AI product survives the next model release only if something valuable remains when the foundation model gets better. Run The Wrapper Test: 8 questions about data, workflow, evals, switching cost, model portability, and buyer pain. Thin wrappers answer with prompts. Durable products answer with evidence.

CRV’s AI wrapper explainer gives the clean definition: a product built on top of an existing foundation model through API calls rather than training a model from scratch. That definition is neutral. A wrapper can be a toy, a cash-flow business, or the first layer of a serious vertical software company.

CRV divides AI wrappers into 4 categories: simple interface wrappers, vertical SaaS wrappers, workflow-embedded wrappers, and agent-based wrappers.

The investor question is not whether the product uses GPT, Claude, Gemini, or an open model. Almost everyone does. The question is whether the product is a restaurant with rented ovens or just a paper menu taped to someone else’s kitchen door. Here is the fixed test.

There’s a runnable version of this test: wrapper-test on GitHub, a free CLI that statically scores any AI product repository against the eight questions below in about a second.

Question 1: what remains if the model ships this feature?

This is the core question. If the answer is “a prompt and an interface,” the product is exposed. If the answer is customer data, workflow history, permissions, integrations, and measured outcomes, the product has assets a model release cannot copy in one launch video.

A good answer sounds concrete: “The provider can copy summarization, but it cannot copy annotated denial letters, broker workflow integrations, and audit history.” A bad answer says the interface is more polished. Polished is nice. It is also a coat of paint.

Ask this before valuation math. If the surviving asset is small, the price should be small too. The build vs buy framework for AI products covers the same pressure from the operator side: rented capability is not the same as owned advantage.

Question 2: what proprietary data improves through use?

Data is durable only when the product creates or captures something competitors cannot scrape. A wrapper that merely sends user text to a model has usage. A stronger product turns usage into labels, corrections, workflow traces, customer-specific rules, or outcome data.

A good answer names the data and how it compounds: “Every reviewed invoice creates a corrected field map tied to that customer’s vendor history.” A bad answer says the product has lots of prompts or stores chat transcripts. Chat transcripts are a junk drawer unless someone turns them into structured learning.

The useful test is replacement. If a customer left tomorrow, what data would make the new vendor worse on day one? If nothing, the moat may be distribution, not product depth.

Question 3: where does it sit in the buyer’s daily work?

Workflow depth means the product lives where the buyer already works, not in a separate tab someone remembers on Fridays. Products embedded in Salesforce, Epic, Outlook, Excel, Procore, a broker portal, or an internal claims system are harder to copy because removal causes operational pain.

A good answer: “Adjusters start and finish the claim inside this queue, and the AI writes back decisions with audit notes.” A bad answer: “Users paste text into the app and copy the result back.” Copy-paste products can grow quickly, but they sit on the counter like a loose appliance. Easy to try. Easy to unplug.

Ask for the screen recording of a normal user session, not the investor demo. The boring workflow tape reveals more than the polished pitch.

Question 4: what eval proves the product is better than the raw model?

The wrapper needs a private scoreboard that compares the full product against the underlying model alone. Without that comparison, nobody knows whether the wrapper adds quality or just packaging. This question turns differentiation into a measured contest.

OpenAI’s eval guidance describes evals as structured tests for accuracy, performance, and reliability despite AI variability. For wrapper diligence, ask for two columns: raw model result and product result on the same real cases.

A good answer: “The product beats the raw model on renewal clauses because it retrieves customer policy history first.” A bad answer: “Customers prefer the experience.” Preference matters after proof. It does not replace proof.

Question 5: can the team swap models without rebuilding?

Model portability is the fire escape. Providers change prices, policies, models, and product boundaries. A startup does not need perfect portability, but it needs enough abstraction and testing that a model change is a project, not an existential event.

OpenAI’s API deprecations page lists model snapshots with scheduled shutdown dates and recommended substitutes. That is normal for a platform. For a thin wrapper with no evals, it is a storm warning.

A good answer: “The eval suite runs on backup models, and the biggest gap is table extraction.” A bad answer: “The product is optimized for one model because it is best.” Best is a weather report, not a roof.

Question 6: what actions happen outside the chat box?

Real products change systems of record. They file tickets, update claims, route approvals, draft filings, reconcile invoices, or trigger reviews with permission checks. A wrapper that only writes text may still be useful, but action depth separates a tool from a talking screen.

A good answer names governed actions: “The system can approve low-risk refunds under a policy limit, route exceptions to a manager, and log every decision.” A bad answer says it can “help with anything.” Anything is not a workflow. It is a fog machine.

This question also reveals security maturity. The moment an AI can take action, prompt injection, access control, and audit logs move from nice-to-have to board-level risk.

Question 7: what would make leaving painful for a customer?

Switching cost should come from useful integration, not contractual handcuffs. The best products become the shelf where customers keep working knowledge: templates, history, approvals, corrections, user roles, reports, and outcome records.

A good answer: “Leaving means rebuilding approval rules and exporting decision history.” A bad answer: “Contracts are annual.” Annual contracts delay churn. They do not create love, data, or better work.

Ask what the product knows about a customer after onboarding that it did not know on day one. If the answer is nothing, the customer has been renting output, not building an operating asset.

Question 8: what specific pain makes the buyer pay?

The last question is commercial, because wrapper risk is not only technical. A wrapper survives if it solves a painful job with a visible budget owner. A generic productivity lift is soft. A reduced denial backlog, faster close process, or fewer compliance review hours is harder to ignore.

A good answer has a buyer, a before-and-after measure, and a budget line. “Claims leaders pay because reviews get faster on this category.” A bad answer says the market is huge because everyone writes emails. A huge market without a sharp pain is a stadium with no seat numbers.

This is also where thin products sometimes pass. A wrapper can be durable if distribution is unfairly strong and the job is narrow enough. But the founder should say that plainly. Hidden thinness is the danger.

How to score the test

Do not average the 8 answers. One fatal answer can be enough. If there is no eval, the team cannot prove differentiation. If there is no data advantage, the product may be a feature. If there is no model fallback, a provider update can turn product planning into damage control.

For an angel check, mark each answer green, yellow, or red in the meeting. For a larger check, hand the greens and yellows to one engineer for a day and ask for a short memo on evals, data flow, model abstraction, and integration depth.

The wrapper label should not kill a deal. It should sharpen the price. A wrapper that compounds data and burrows into daily work can become a serious company. A wrapper that depends on one prompt and one provider release calendar is a trade, not a foundation.