Field notes on
production AI.
What a senior engineering team is reading, shipping, and arguing about — written for people who build AI in systems that can't fail.
AI startup red flags: the technical ones investors miss
Eight technical AI startup red flags non-technical investors can catch by ear in a pitch, with the follow-up question that confirms each one.
AI due diligence vs software due diligence
Why AI deals break standard software diligence: 4 gaps code scans miss, from eval quality to wrapper depth, before investors write a check this month.
How to evaluate an AI startup without a technical background
A no-code investor guide with 6 questions, good and bad answer patterns, and when to bring one engineer in before a term sheet or check for AI seed deals.
The Wrapper Test: will this AI product survive the next model release?
An 8-question Wrapper Test for investors: how to tell whether an AI product has data, workflow depth, evals, and customers a model release cannot copy.
AI agent platforms compared, before you commit
Compare AWS Bedrock AgentCore, OpenAI, Claude with MCP, LangGraph, and custom builds on lock-in, cost, and control before OpenAI's August 26, 2026 shutdown.
Build vs buy for AI products
A build vs buy framework for AI products: why 76% of enterprise AI is now bought, when building wins anyway, and the switching-cost math founders skip.
Fractional CTO vs dev agency vs technical co-founder
Fractional CTOs bill $150 to $500 an hour, US agencies charge $100 to $149, and a co-founder search can take a year. What each option really costs you.
How much does it cost to build an AI product in 2026
A verified 2026 cost breakdown for building an AI product: salaries, per-token API prices, data work, and why regulated systems cost 3 to 10 times more.
How to evaluate an AI startup's technology before investing
Five questions that reveal whether an AI startup's technology is real before you invest, and what honest answers sound like when $445 million is at stake.
Predicting AI agent performance without the huge cost
A new framework, PACE, predicts complex AI agent performance with over 80% accuracy, cutting evaluation costs to less than 1% of traditional methods.
Questions to ask an AI vendor about security
Seven security questions to ask any AI vendor before you sign, with the good answer and the dodge for each, from training data to SOC 2 and incident history.
A technical due diligence checklist for AI startups
The 7 checks that matter before wiring money into an AI startup: what to ask about evals, data rights, unit economics, and what a bad answer sounds like.
When not to use AI agents
Gartner expects over 40% of agentic AI projects to be canceled by 2027. A practical four-question test for spotting tasks where an agent is the wrong tool.
How AI learns to check its own work using visual cues
A new training method, Visually Grounded Reinforcement Reflection Learning (VRRL), helps AI models improve their accuracy on unseen visual tasks by 25% by learning to reflect on their mistakes.
Keeping your AI models secure without slowing them down
NVIDIA Confidential Computing offers a way to protect your AI models and data during use, maintaining up to 98% of performance.
Securing AI without slowing it down: NVIDIA's confidential computing
NVIDIA Confidential Computing offers hardware-level AI security, maintaining up to 98% of native inference performance during active use.
Why the AI that writes code needs to think more, not just have more tools
A new study shows that AI agents building software are more reliable when they focus on reasoning, not just accessing more tools. Results are up 89%.
How to make AI assistants remember the right things
AI assistants struggle with memory recall, but a new technique improves accuracy by 24% by filtering based on specific criteria.
How synthetic data and fine-tuning improve vision AI agent accuracy
Many businesses struggle to train AI that sees. This post explains how synthetic data and fine-tuning can boost vision AI agent accuracy by 95% or more, even for rare events.
Automating healthcare claims with AI agents and Amazon Bedrock
New AI approaches can cut manual healthcare claims processing time by 37%, reducing errors and speeding up payments.
How AI agents are unmasking 'anonymous' location data
New research shows AI agents can re-identify 72% of individuals from 'anonymous' location data, raising urgent privacy concerns.
How AI agents securely use your company's data with a modern data mesh
AI agents go beyond chatbots, performing tasks by accessing company data. Learn how a modern data mesh on AWS provides a secure foundation, potentially cutting vector storage costs by 90%.
How to safely run AI agents in your business
AI agents can boost productivity by 37%, but they need strict governance to protect sensitive data and systems. Learn how to control them.
Mapping Europe's AI workforce opportunity
A new report estimates roughly 37% of tasks in European jobs could be affected by AI, creating both challenges and fresh opportunities for businesses.
Securing AI analytics: Why trusting your AI agent with data access is a mistake
Discover how a three-layer security architecture prevents data leaks in multi-tenant AI analytics systems, protecting sensitive business information for over 300 restaurant brands.
Stripe's lesson: Why focused AI agents excel in compliance
Stripe handles $1.4 trillion in payments annually. Learn how they use AI agents to cut compliance review time by 26% while keeping humans in control.
What can AI agents actually do today?
A new index reveals where 300 builders see AI agents delivering real value, with confidence scores for 101 tasks.
When AI checks its own homework: a surprising look at how models evaluate answers
A new study reveals that large language models often struggle to reliably evaluate answers, performing worse than their generation capabilities in 3 out of 4 benchmarks.
Why a smart AI needs a self-evolving map of the world to get things done
Discover how a new AI approach, WorldEvolver, enhances agent reliability by revising its internal world model, leading to up to 37% higher success rates in complex tasks.
Why AI sometimes misreads numbers in your spreadsheets
Large language models can make 'data referencing errors' when reading tables, but a new approach improves accuracy by up to 12.0%.
Why your AI agents need an immune system
Autonomous AI agents can be hijacked at runtime. A new defense, the Agent-Native Immune System, offers internal protection, reducing operational risks by up to 37%.
Why your AI assistant might be sharing too much private data
New research reveals that 9 widely used AI agents often overshare private business data, even when completing tasks successfully, posing a significant risk.
Why your AI's vision might be failing even when it seems smart
New research shows AI models can miss critical visual details, revealing an 8% perception gap between open and proprietary systems, despite high benchmark scores.
How insurance brokerages are using AI to get work done faster
Insurance brokerages are saving about 10 hours per user per week with AI that understands their specific business needs, not just general chat.
How AI agents are doing more than just talking
AI agents represent a major shift, moving beyond chatbots to autonomously perform complex tasks, potentially reducing operational time by 37% or more.
How AI agents in Excel can automate complex finance workflows
Microsoft's Copilot in Excel uses AI agents and custom skills to automate multi-step financial tasks, connecting to over 6 financial data providers.
How NVIDIA's Nemotron 3 Ultra helps AI agents get more done
NVIDIA's Nemotron 3 Ultra, a new 550B-parameter AI model, helps AI agents complete complex tasks 5x faster and with 30% lower cost.
How a small AI model can speed up building bigger, smarter ones
A new research framework called Knowledge Cascade can cut the computational cost of developing large AI models by up to 37%, making advanced AI more accessible.
How AI can now use your computer to get things done
A new AI capability, 'computer use' in Gemini 3.5 Flash, lets AI agents perform tasks across applications, potentially cutting manual work by 37%.
How an AI agent can handle healthcare appointments by voice
Learn how an AI voice agent, powered by Amazon Nova 2 Sonic, can reduce healthcare appointment no-show rates by 5-30% by automating patient conversations.
How banks can find and hide sensitive customer data in millions of documents
Huntington Bank used AWS services to automatically find and redact sensitive data from over 400 million documents in just months, not years.
Why your AI investments might be sitting idle: understanding GPU usage
Many companies underutilize their expensive AI hardware, with up to 75% of GPUs sitting idle. A new monitoring tool provides real-time visibility.
How AI agents pay for things on their own
AI agents are starting to pay tiny amounts per task instead of running on a flat subscription. A look at how Ampersend and Amazon Bedrock make that work safely.
Why warmer coolant makes AI data centers more efficient
NVIDIA's new AI servers use 45°C liquid cooling, cutting data center energy use by up to 40% and saving a 50-megawatt facility over $4 million annually.
Why an AI agent needs web search to stay current
Discover how Web Search on Amazon Bedrock AgentCore updates AI agents with fresh information in minutes, enabling better real-time business decisions.
Amazon Bedrock AgentCore harness: from idea to a working agent
Amazon released a tool that handles the setup work behind AI agents. Here is what an agent is, what the harness does, and why the timing matters.
How ChatGPT Enterprise now shows what your team spends
OpenAI added usage analytics and spend controls to ChatGPT Enterprise. Admins can now see who uses what and set monthly limits before costs climb.
Agentic resource discovery: letting AI agents find their own tools
AI agents can usually only use tools a developer set up in advance. A new shared standard lets an agent search for the right tool on its own, mid-task.
AI agents that write their own code to answer questions about company data
A new paper shows AI that writes and fixes its own code to read a company's data, tested on seven SQL benchmarks, matching or beating the best results.
AI is helping attackers move faster, and what it means for you
The same kind of AI behind ChatGPT now helps online attackers work faster. A Microsoft post explains how, and what a normal person can do to stay ahead.
Can AI agents be trusted in drug discovery? What TxBench-PP found
A new test gave 11 AI models real drug-lab data. The best one got the right answer 59.3 percent of the time, and none of them could be trusted on their own.
LifeSciBench: how OpenAI measures AI on real science
A plain-English look at LifeSciBench, OpenAI's new test that checks how well AI handles real life science research, and what its low scores tell us.
Why AI agents give wrong answers about your company data
An AI agent is only as smart as the context it can reach. AWS Context maps how a company's scattered data connects so agents stop guessing.
What the Amazon Bedrock InvokeGuardrailChecks API means for running AI agents safely
Amazon released a new safety tool on 16 June 2026 that checks an AI for risk at every step of its work, not just at the start, and gives each check a score instead of blocking on its own.
Can AI models remember what they saw? What RNG-Bench found
Researchers built a test to check whether 7 AI models can act on facts that are no longer on the screen. The best one got 62.3 percent of card pairs right, and bigger boards broke them fast.
How companies actually get value from AI
A plain guide to what a company has to figure out before AI pays off, built around three questions every business is asking about it right now.
Why AI apps can be slow to start, and how caching speeds them up
When an AI app gets busy, it has to wake up new computers to keep up. Those computers download a big file first, and that wait is what caching now cuts.
CEO-Bench: can AI agents run a company for 500 days?
A Princeton test puts AI agents in charge of a fake software company for 500 days. Most went broke. Two finished rich. Here is what that tells you.
Why AI agents fail in production and how failure detection fixes it
When an AI agent gets something wrong, the mistake usually hides several steps back. A new AWS tool reads the agent's own logs and tells you which step broke and why.
Teaching AI to use a computer by letting it practice
Most AI only chats. A new report shows how teaching an AI to click and type its way through a real desktop made it far better at finishing tasks.
How Rocket Close cut title work with an AI that takes actions
A real company put an AI agent on its slowest, most paperwork-heavy task and cut contact center calls and emails by 30 percent. Here is what it did, in plain terms.
How to tell if an AI agent is actually working
A right answer can hide a broken AI agent. A new AWS toolkit checks the steps, not just the final reply, and one test scored faithfulness at 32 percent.
Turning piles of documents into clean data without typing it in
How AI reads scanned PDFs and pulls out the useful facts on its own. A plain walk through fast one-at-a-time versus cheaper overnight processing for beginners.
The price of anarchy: why serving AI cheaply is a traffic problem
A new paper studies what happens when an AI service gets crowded. Above a certain load, costs grow more than 280 times. Here is why, in plain words.
No posts under this topic yet.