The demo arrives at minute twelve of the pitch. The founder types one sentence, the screen fills with a working app, and the whole thing takes 40 seconds. Nobody in the room says anything for a moment. That silence is worth millions, and founders know it.

Builder.ai produced that silence for years. The London startup sold an AI that supposedly assembled software the way you’d order a pizza, and it raised more than $445 million from backers including Microsoft and the Qatar Investment Authority before entering insolvency in May 2025. Workers later told Rest of World that much of the work marketed as AI had been done by human developers in India and Ukraine, and the company had already restated its 2023 revenue down to $140 million after earlier projections collapsed. The demos were persuasive right up until the accountants arrived.

A written technical due-diligence checklist covers the documents: architecture, code, contracts, infrastructure bills. This piece covers something different. It covers the 90 minutes you spend in a room with the founding team, and the five questions that separate real technology from theater. You don’t need to write code to ask any of them. You need to know what a bluff sounds like, and what honesty sounds like, because they are surprisingly easy to tell apart once you know the pattern.

Question one: can I drive?

When you buy a used car, the salesman doesn’t take the test drive for you. Yet in AI pitches, investors routinely watch the founder drive for twenty minutes and call it evaluation.

Even Google plays this game. In December 2023 it released a Gemini demo video that appeared to show the model responding fluidly to live voice and video, and TechCrunch reported within a day that the interaction was stitched together from still images and text prompts, with latency edited out. If a trillion-dollar company polishes its demos that hard, assume a seed-stage founder does too.

So take the keyboard. Type your own inputs, ideally something from your world: a portfolio company’s support tickets, or just a contract you happen to have handy. The bluff has a recognizable shape. The founder keeps control of the machine, steers you back to prepared examples, and answers unscripted requests with “great idea, let me follow up on that.” An honest team does the opposite. They hand you the laptop, they seem mildly excited to see what you’ll try, and when the product stumbles they narrate the failure. “Yeah, it’s weak on tables, that’s our current sprint” is one of the most bullish sentences you can hear in a pitch.

Question two: what here is actually yours?

“We built a proprietary model” is the most abused sentence in AI fundraising. Sometimes it’s true. Usually it means the team took an existing model built by someone else and fine-tuned it, which is the practice of adjusting a finished model with a relatively small set of your own examples. Think of a chef claiming a secret sauce that turns out to be a store-bought base with extra paprika. The dish might still taste great. But you shouldn’t pay recipe-inventor prices for it.

The economics make the gap concrete. Training a frontier model from scratch costs tens of millions of dollars in computing power alone. Fine-tuning GPT-4o through OpenAI’s public service costs $25 per million training tokens, so a serious fine-tuning run frequently lands under a few thousand dollars. Those two activities differ in cost by roughly a factor of ten thousand, and only one of them justifies a deep-tech valuation.

And here’s the nuance: being a wrapper is not a crime. Some excellent businesses are thin layers over GPT-5 or Claude, with the real value in workflow and distribution and accumulated customer data. What matters is whether the founder describes it honestly. A bluffing founder gets vague when you ask what happens if OpenAI ships their feature next quarter. An honest one says something like “the model is rented, the data pipeline and the eight months of customer feedback loops are ours,” and can point to which specific components would survive a model swap. If the pitch involves handling sensitive customer data, this is also the moment to fold in the security questions you’d ask any AI vendor, because rented models mean data flows through third parties.

Question three: show me the eval suite

This is the single highest-signal request in AI diligence, and most investors have never made it. An eval suite (short for evaluation suite) is the private battery of test cases a team runs against its own product: hundreds or thousands of real inputs with known correct answers, scored automatically before every release. It’s the practice exam the team grades itself with. A company that has one can tell you, with numbers, whether last month’s changes made the product better or worse. A company that doesn’t have one is guessing, no matter how confident the founder sounds.

Public benchmark scores are not a substitute, because benchmarks leak. Researchers at Scale AI rebuilt a popular math benchmark from scratch with fresh questions of identical difficulty, and found that some model families scored up to 13% worse on the new version. The models had partially memorized the old test rather than learned the skill. A student who aces the exact practice test he’s seen before tells you nothing about the final exam. So when a deck quotes benchmark numbers, ask a quiet follow-up: are those scores for your product on your customers’ tasks, or for the underlying model on a public test it may have memorized?

The honest answer to “show me your evals” is refreshingly boring. Someone opens a dashboard, you see pass rates over time, and you see failures logged and categorized. There will be regressions on the chart. Real charts have dips. The bluff is a founder who quotes GPT-5’s public scores as if they were the product’s, or promises to send materials that arrive as a marketing PDF.

Question four: what did you try that didn’t work?

Real engineering leaves scar tissue. A team that has genuinely wrestled with AI in production has a graveyard of dead ends, and honest founders talk about it with specific dates and specific pain. “We spent six weeks on an agent architecture and ripped it out in March because it kept looping” is what truth sounds like. So is a founder who can name the customer segment where the product currently fails.

Silence here is the tell. If every technical decision in the company’s history was apparently correct on the first try, either the team hasn’t pushed the technology anywhere hard, or you’re being managed. Follow up by asking what inputs break the product today. Every AI system has them, because these systems fail in strange, unpredicted ways rather than cleanly. The teams that thrive treat failure modes as inventory to be tracked, something covered in more depth in how much it really costs to build an AI product in 2026, where post-launch fixing routinely outweighs the initial build.

Question five: who built the core, and are they still here?

Code can be audited. Conviction can’t, except one way: watch what the engineers do with their own careers. Startup engineers typically hold stock options on four-year vesting schedules with a one-year cliff, which means walking away early burns real money. A senior engineer who quits eight months before a funding round is making a priced bet against the company, with better information than any outside investor will ever have.

So ask directly. Who wrote the core system? Are those people still employed? Has any technical leader left in the past year, and why? Then make the request that bluffing founders hate: an hour alone with the lead engineer, no founder in the room. Honest companies agree easily, and the engineer talks about problems with the unguarded enthusiasm of someone who likes their work. Evasive companies route every technical question through the CEO, and at Builder.ai the gap between the sales story and what the engineers privately knew stayed hidden for years precisely because nobody outside kept asking the people who did the work.

None of these five questions requires you to read a line of code. They require 90 minutes and a willingness to be politely stubborn. Bringing one experienced engineer along for the day helps too, and costs a few thousand dollars against a check that might be seven figures. The pattern underneath all five is the same. Real technology survives contact with strangers. Theater needs the founder’s hands on the keyboard, and the moment you reach for it yourself, you’ll know which one you’re looking at.