Thursday’s wire is waiting, and the founder just showed you an AI demo you cannot inspect. Use the Across-the-Table Test: six plain questions that expose measurement, ownership, data rights, cost, resilience, and engineering judgment. You are not trying to code-review the company. You are listening for proof.
This page is the no-code entry point. If the answers survive it, the deeper reads are the technical due diligence checklist for AI startups and the more hands-on guide to evaluating an AI startup’s technology before investing. Start here because it keeps the first meeting human. One table. One founder. One investor asking questions a real engineer would ask herself before touching the repository.
The National Institute of Standards and Technology names 7 characteristics of trustworthy AI, including validity, safety, security, accountability, explainability, privacy, and fairness in its AI Risk Management Framework.
That sounds abstract until money is on the table. Think of it like buying a restaurant you cannot cook in. You do not need to know the sauce recipe first. You need to know whether the kitchen has health records, supplier invoices, repeat customers, and one person who can open the oven without guessing.
The Across-the-Table Test
The Across-the-Table Test is the set of questions a practical engineer would ask while the pitch deck is open: what is measured, what is owned, what breaks, what costs money, what data is allowed, and who can fix it. Good answers contain receipts, not adjectives.
The six questions are deliberately boring. Boring is good in diligence. A founder who can answer them without performing is usually close to reality. A founder who keeps returning to the demo is asking you to judge a kitchen by one plated dish.
Use the same script for every AI deal. The consistency matters. If every company gets a different interview, you end up comparing personalities. If every company gets the same six questions, you start comparing evidence.
Question 1: what score do you watch every week?
AI quality needs a scoreboard, because demos are too easy to polish. Ask the founder which eval, meaning a structured test of model output, the team watches every week and how the score moved after the last product change. The answer should sound like operations, not theater.
OpenAI’s eval guide says evals are structured tests for accuracy, performance, and reliability despite AI variability. That is the plain idea. If the product summarizes loan documents, there should be a pile of old loan documents with known correct answers. If it handles support tickets, there should be real tickets with graded outcomes.
A good answer: “The billing-case eval dipped after the April routing change, and the misses are tagged by failure type.” A bad answer: “Users love it and the model is getting better every month.” One is a thermometer. The other is a hand on the forehead.
Question 2: what survives if the model provider ships this?
Most AI startups rent intelligence from OpenAI, Anthropic, Google, or an open model host. Renting is normal. The investor question is what the startup owns besides rent payments: customer data, workflow depth, distribution, compliance paperwork, and a test harness that makes models replaceable.
Ask the founder: “If your model provider ships this feature next quarter, what still belongs to you?” A good answer names assets the provider cannot copy from an API log. The team has proprietary data from customer work, integrations into tools buyers use each morning, or domain rules that took months to encode.
A bad answer leans on prompt secrecy. Prompts are like a sticky note on a kitchen wall. Useful, yes. A business, no. If the only moat is a clever instruction written into a text box, price the company like a feature, not a platform.
Question 3: what breaks today?
Real AI products have known weak spots. The best founders can name them quickly because they argue with those failures every week. Ask what input currently breaks the product, which customer segment is not ready, and what the team decided not to automate in 2026.
A good answer has scars. “It fails on scanned tables, so those cases route to human review.” Or: “It works in English contracts but not German renewals yet.” A bad answer says the product is strong across all document types, all industries, and all languages. No tool is a Swiss Army knife, a dishwasher, and a bridge crane at the same time.
This question is less technical than it sounds. You are checking whether the founder has a map with dangerous roads marked in red. A founder with no red roads may be brave. More likely, nobody has driven there.
Question 4: where did the data come from?
Data provenance is the chain of custody for the examples, documents, recordings, or customer records used by the AI. Ask where the data came from, who gave permission, and whether that permission covers the specific AI use being pitched. A clean answer sounds like a filing cabinet.
Good: “Customers upload their own documents, the contract allows processing for this purpose, and training data is separated from live customer data.” Bad: “It is public data.” Publicly visible is not the same as legally usable. A library book is public enough to read in the building. It is not public enough to photocopy into your own textbook.
If the startup claims it trained or fine-tuned a model, ask for the source list. Not the raw data. The source list. The distinction matters because a responsible team can prove custody without handing you customer records.
Question 5: what does one normal action cost?
AI products have a meter running inside them. Every model call, search, document parse, and retry costs money, even when the interface looks like a simple button. Ask for the cost of one normal customer action and the margin after that action.
A good answer sounds like a receipt: “A normal claims review uses document extraction, model calls, and a tracked cloud cost before storage.” A bad answer says token prices keep falling. That may be true in general, but it is not a margin model. A cafe owner cannot answer food cost with “wheat gets cheaper over time.”
The exact number matters less than whether the founder has one. If the team cannot price one unit, it cannot forecast gross margin at customer scale.
Question 6: who can fix it without the founder?
Small startups are allowed to be fragile. They are not allowed to be unaware. Ask who can deploy a fix if the technical founder is unavailable, where the runbook lives, and how the team rolls back a bad release on a Friday.
A good answer names another person and a process: “Mira can deploy, the rollback is in GitHub Actions, and incidents are written up in Linear.” A bad answer is a smile followed by “that has never happened.” Software that has never had an incident is usually software that has not had enough users.
This is the bus-factor question without the bluntness. You are not demanding a 30-person engineering org. You are asking whether knowledge lives in a system or in one head.
When to bring one engineer for a day
Bring an engineer after the first pass, not before it. The right moment is when the founder answered the six questions well, the check is still attractive, and one technical claim controls the valuation. One day is enough for the first independent read.
The engineer should not boil the ocean. Give her four jobs: read the architecture diagram, inspect the eval setup, trace customer data through the system, and ask how a model swap would work. That is a day of work, not a month of consulting. The output should be a short risk memo, not a full acquisition audit.
The Angel Capital Association due diligence guide cites one New York Angels process that expects around 30 conversations during reference work.
That number is a useful reminder. Diligence is not one clever question. It is repeated contact with reality. For an AI deal, the first contact is this test. The second is a focused engineer review. The third is customer reference work that asks whether the product keeps working after the demo laptop is closed.
What good and bad answers sound like
Good answers are specific before they are impressive. They contain dates, tools, counts, contracts, logs, and the occasional admission of weakness. Bad answers are smooth. They contain phrases like “very accurate,” “fully proprietary,” “enterprise-ready,” and “public data” where a document should be.
One sentence is especially bullish: “That is not built yet.” It sounds weak, but only when it is alone. The stronger version is, “That is not built yet, here is the owner, here is the date, and here is the temporary control.” That is how builders talk.
Your job is not to become the CTO in the room. Your job is to make the founder move from story to evidence. If the evidence holds, go deeper. If it does not, you just saved yourself from paying for a magic trick.