AI startup red flags show up in how founders answer plain questions, not only in code. Listen for missing failures, vague data claims, single-provider dependence, laptop demos, benchmark theater, and blank scale answers. Each one points to engineering risk you can catch before hiring a technical reviewer.

Use this as the room script before the deeper technical due diligence checklist for AI startups. If you do not have a technical background, pair it with how to evaluate an AI startup without a technical background. The goal is not to sound technical. The goal is to make a polished pitch touch reality.

1. The founder cannot name a failed eval

A founder who cannot name a failed eval is not proving the product is reliable. They are revealing that the team does not inspect failure in a repeatable way. Serious AI teams know the misses by category, because those misses decide what ships, what waits, and what needs human review.

What it sounds like: “We test constantly,” followed by no remembered failure, no failure log, and no example that made the team change the product.

Why it matters: an eval is the product’s exam. If no one can name a failed answer, the team may be learning from demos, customer complaints, and vibes instead of a scoreboard.

Follow-up question: “Show me one eval case the product failed recently, what caused it, and what you changed because of it?“

2. Proprietary data has no compounding story

“Our data is proprietary” is weak until the founder can explain how it gets better through use. Data only matters when the product creates labels, corrections, workflow traces, customer rules, or outcome records competitors cannot recreate by scraping public pages or buying a generic dataset.

What it sounds like: “We have unique data,” but the answer never reaches source, permission, structure, freshness, or how one customer’s work improves the next result.

Why it matters: static data is an ingredient. Compounding data is a kitchen system. The first can be copied or purchased. The second can make the product harder to replace over time.

Follow-up question: “What new data does each customer create, and how does that make the product better next month?”

This is where ordinary code review misses the point. AI due diligence vs software due diligence is mostly the difference between inspecting a repository and inspecting the evidence around behavior.

3. One-provider dependency is sold as a feature

Depending on one model provider is not automatically bad, especially early. The red flag is when dependence is described as strategy. If the product works only because one provider currently has one model, one API behavior, and one pricing shape, the startup has a hidden landlord.

What it sounds like: “We are all in on this provider because it is the best,” with no backup model, no migration notes, and no eval comparison.

Why it matters: providers change models, prices, policies, and product boundaries. A startup that cannot switch is not just using a platform. It is letting the platform write part of its roadmap.

Follow-up question: “If your primary provider changed pricing, policy, or shipped your main workflow, what would still work on a backup model?”

This is the same pressure behind The Wrapper Test: what survives when the model provider moves closer to the customer?

4. No one has run this in production

An AI prototype can look mature before anyone has operated it under real customer pressure. Production experience means someone has owned incidents, retries, latency, cost spikes, data mistakes, customer permissions, and rollback. Without that scar tissue, the team may be learning basic operations with your money.

What it sounds like: “We are technical founders,” but nobody has shipped and maintained a customer-facing AI, data, or workflow product after users depended on it every day.

Why it matters: production is where demos meet weather. The model times out, the upload fails, the customer pastes strange data, the provider rate-limits a queue, and someone has to know what happens next.

Follow-up question: “Who on this team has owned a production system with real users, and what broke while they owned it?“

5. The roadmap is all model-provider features

If the roadmap sounds like future model release notes, the company may be waiting for its supplier to build the product. Longer context, better tool use, lower latency, lower cost, and stronger reasoning may help everyone. They are not a company-specific plan unless they unlock owned workflow or data.

What it sounds like: “As models improve, we will handle more cases,” with no customer workflow, integration, data asset, compliance step, or distribution advantage that becomes stronger.

Why it matters: provider improvements are shared wind. They lift competitors too. A startup needs something that becomes more valuable when models improve, not a plan that disappears when models improve.

Follow-up question: “Which roadmap item still matters if the next model release includes the visible feature you are building?“

6. The demo only runs on the founder’s laptop

A laptop-only demo means the pitch may be showing a performance, not a deployable product. The question is not whether the founder can make it work. The question is whether the company can run it repeatably, with normal setup, customer-like data, access control, logs, and monitoring.

What it sounds like: “The hosted version is being cleaned up,” “staging has some issues,” or “I will drive because the setup is a little sensitive.”

Why it matters: a local demo can hide missing authentication, brittle data paths, manual fixes, hardcoded examples, weak observability, and deployment steps nobody else understands. Those are not polish issues. They are operating risk.

Follow-up question: “Can someone other than you open a fresh environment and run the same workflow with a customer-like account?“

7. Metrics are benchmark scores, not production numbers

Benchmark scores are model facts, not company facts. They may explain why the team chose a model, but they do not prove that the startup’s product works for customers. Investors need production metrics tied to the workflow the customer actually pays to improve.

What it sounds like: “The underlying model performs well on public benchmarks,” followed by no task completion rate, review rate, failure category, latency, cost, or customer outcome.

Why it matters: public benchmarks test controlled tasks. Production numbers test the product’s real job. A claims product, finance copilot, or legal assistant can use an impressive model and still fail where the buyer needs precision.

Follow-up question: “What production metric do you watch every week, and how did it move after the last release?“

8. The answer to “what breaks at 10x scale?” is blank

The answer to what breaks at 10x scale should be specific. A strong founder can name the first bottleneck before it arrives. A weak answer treats scale as a cloud setting, even though AI scale hits model cost, queues, rate limits, review capacity, data handling, and support.

What it sounds like: “Nothing major, the cloud scales,” or a long pause before a generic answer about hiring more engineers.

Why it matters: growth magnifies weak systems. If the team cannot predict the first constraint, it probably has not modeled the product as an operating system with costs, limits, people, and failure paths.

Follow-up question: “At 10x current usage, which limit hits first, how do you know, and what is the mitigation plan?”

How to read the pattern

You are not trying to prove the startup is bad. You are checking whether the team knows reality. Strong founders answer with evidence, including ugly evidence. Weak founders return to the demo, the benchmark, the provider, or the vision when the question asks for operations.

One red flag does not kill every deal. Seed companies are incomplete by design. The useful distinction is awareness. “We do not have that yet, and here is the owner and date” is very different from “that will not be a problem.” The first is a plan. The second is a bet with your check underneath it.

Use the questions in the meeting. Mark each answer green, yellow, or red. If most answers are green, bring one engineer in for a focused review of evals, data flow, model fallback, and deployment. If most answers are red, you did not need code access. You heard the engineering rot through the polish.