Writing

Field notes on
production AI.

What a senior engineering team is reading, shipping, and arguing about — written for people who build AI in systems that can't fail.

Subscribe via RSS
An investor listening for technical red flags behind a polished AI startup demo
Jul 9, 2026 · 7 min

AI startup red flags: the technical ones investors miss

Eight technical AI startup red flags non-technical investors can catch by ear in a pitch, with the follow-up question that confirms each one.

investing / AI due diligence
Two diligence checklists, one for software code and one for AI behavior
Jul 6, 2026 · 8 min

AI due diligence vs software due diligence

Why AI deals break standard software diligence: 4 gaps code scans miss, from eval quality to wrapper depth, before investors write a check this month.

AI due diligence / software diligence
An investor reviewing an AI startup pitch with a checklist
Jul 6, 2026 · 8 min

How to evaluate an AI startup without a technical background

A no-code investor guide with 6 questions, good and bad answer patterns, and when to bring one engineer in before a term sheet or check for AI seed deals.

AI due diligence / investing
A thin AI app layer tested against a larger foundation model release
Jul 6, 2026 · 7 min

The Wrapper Test: will this AI product survive the next model release?

An 8-question Wrapper Test for investors: how to tell whether an AI product has data, workflow depth, evals, and customers a model release cannot copy.

AI wrappers / AI investing
Five doors side by side, each a different AI agent platform choice
Jul 4, 2026 · 7 min

AI agent platforms compared, before you commit

Compare AWS Bedrock AgentCore, OpenAI, Claude with MCP, LangGraph, and custom builds on lock-in, cost, and control before OpenAI's August 26, 2026 shutdown.

AI agents / platforms
A fork in the road between buying an AI platform and building one in-house
Jul 4, 2026 · 7 min

Build vs buy for AI products

A build vs buy framework for AI products: why 76% of enterprise AI is now bought, when building wins anyway, and the switching-cost math founders skip.

build vs buy / AI strategy
Three diverging paths representing three ways to get an AI product built
Jul 4, 2026 · 7 min

Fractional CTO vs dev agency vs technical co-founder

Fractional CTOs bill $150 to $500 an hour, US agencies charge $100 to $149, and a co-founder search can take a year. What each option really costs you.

founders / hiring
Stacked cost bars comparing an AI MVP budget to a regulated production system budget
Jul 4, 2026 · 7 min

How much does it cost to build an AI product in 2026

A verified 2026 cost breakdown for building an AI product: salaries, per-token API prices, data work, and why regulated systems cost 3 to 10 times more.

AI costs / founders
A magnifying glass held over a polished product demo screen
Jul 4, 2026 · 8 min

How to evaluate an AI startup's technology before investing

Five questions that reveal whether an AI startup's technology is real before you invest, and what honest answers sound like when $445 million is at stake.

investing / AI due diligence
Diagram showing how a proxy benchmark predicts complex AI agent performance.
Jul 4, 2026 · 8 min

Predicting AI agent performance without the huge cost

A new framework, PACE, predicts complex AI agent performance with over 80% accuracy, cutting evaluation costs to less than 1% of traditional methods.

AI agents / LLM evaluation
A checklist beside a vendor contract and a padlock
Jul 4, 2026 · 8 min

Questions to ask an AI vendor about security

Seven security questions to ask any AI vendor before you sign, with the good answer and the dodge for each, from training data to SOC 2 and incident history.

AI security / procurement
A numbered checklist beside a stack of investment documents
Jul 4, 2026 · 8 min

A technical due diligence checklist for AI startups

The 7 checks that matter before wiring money into an AI startup: what to ask about evals, data rights, unit economics, and what a bad answer sounds like.

AI due diligence / investing
A fork in the road with one path marked automation and one marked human review
Jul 4, 2026 · 8 min

When not to use AI agents

Gartner expects over 40% of agentic AI projects to be canceled by 2027. A practical four-question test for spotting tasks where an agent is the wrong tool.

AI agents / risk
Bar chart comparing AI model accuracy with and without visual self-reflection.
Jul 2, 2026 · 8 min

How AI learns to check its own work using visual cues

A new training method, Visually Grounded Reinforcement Reflection Learning (VRRL), helps AI models improve their accuracy on unseen visual tasks by 25% by learning to reflect on their mistakes.

AI agents / machine learning
Diagram illustrating NVIDIA Confidential Computing protecting AI models during inference.
Jul 2, 2026 · 7 min

Keeping your AI models secure without slowing them down

NVIDIA Confidential Computing offers a way to protect your AI models and data during use, maintaining up to 98% of performance.

AI Security / Confidential Computing
Bar chart showing minimal performance overhead of NVIDIA Confidential Computing.
Jul 2, 2026 · 9 min

Securing AI without slowing it down: NVIDIA's confidential computing

NVIDIA Confidential Computing offers hardware-level AI security, maintaining up to 98% of native inference performance during active use.

AI security / Confidential Computing
An AI agent considering different paths to write code
Jul 2, 2026 · 7 min

Why the AI that writes code needs to think more, not just have more tools

A new study shows that AI agents building software are more reliable when they focus on reasoning, not just accessing more tools. Results are up 89%.

AI agents / software development
Diagram showing AI assistant memory filtering
Jul 1, 2026 · 10 min

How to make AI assistants remember the right things

AI assistants struggle with memory recall, but a new technique improves accuracy by 24% by filtering based on specific criteria.

AI agents / memory management
Flow diagram showing how synthetic data and fine-tuning work together to improve AI model accuracy for vision agents.
Jun 30, 2026 · 8 min

How synthetic data and fine-tuning improve vision AI agent accuracy

Many businesses struggle to train AI that sees. This post explains how synthetic data and fine-tuning can boost vision AI agent accuracy by 95% or more, even for rare events.

AI agents / Vision AI
Diagram showing an AI agent processing a healthcare claim form using Amazon Bedrock and AWS HealthLake.
Jun 29, 2026 · 9 min

Automating healthcare claims with AI agents and Amazon Bedrock

New AI approaches can cut manual healthcare claims processing time by 37%, reducing errors and speeding up payments.

AI agents / healthcare AI
Diagram showing an AI agent connecting dots from public data to location trails
Jun 29, 2026 · 8 min

How AI agents are unmasking 'anonymous' location data

New research shows AI agents can re-identify 72% of individuals from 'anonymous' location data, raising urgent privacy concerns.

AI agents / data privacy
Diagram showing an AI agent interacting with various governed data sources through a secure gateway.
Jun 29, 2026 · 9 min

How AI agents securely use your company's data with a modern data mesh

AI agents go beyond chatbots, performing tasks by accessing company data. Learn how a modern data mesh on AWS provides a secure foundation, potentially cutting vector storage costs by 90%.

AI agents / data mesh
Diagram showing a user device connecting to a managed AI agent workspace within an enterprise AI factory.
Jun 29, 2026 · 11 min

How to safely run AI agents in your business

AI agents can boost productivity by 37%, but they need strict governance to protect sensitive data and systems. Learn how to control them.

AI agents / enterprise AI
Diagram showing AI's impact on job tasks across different European sectors.
Jun 29, 2026 · 9 min

Mapping Europe's AI workforce opportunity

A new report estimates roughly 37% of tasks in European jobs could be affected by AI, creating both challenges and fresh opportunities for businesses.

AI workforce / European economy
A layered architecture diagram showing security controls for an AI analytics agent.
Jun 29, 2026 · 9 min

Securing AI analytics: Why trusting your AI agent with data access is a mistake

Discover how a three-layer security architecture prevents data leaks in multi-tenant AI analytics systems, protecting sensitive business information for over 300 restaurant brands.

AI agents / data security
Diagram showing an AI agent system for compliance with human oversight
Jun 29, 2026 · 8 min

Stripe's lesson: Why focused AI agents excel in compliance

Stripe handles $1.4 trillion in payments annually. Learn how they use AI agents to cut compliance review time by 26% while keeping humans in control.

AI agents / Financial Compliance
Illustration of AI agents working on tasks
Jun 29, 2026 · 6 min

What can AI agents actually do today?

A new index reveals where 300 builders see AI agents delivering real value, with confidence scores for 101 tasks.

agentic ai / AI
AI chatbot checking a document with a red pen, looking confused
Jun 29, 2026 · 7 min

When AI checks its own homework: a surprising look at how models evaluate answers

A new study reveals that large language models often struggle to reliably evaluate answers, performing worse than their generation capabilities in 3 out of 4 benchmarks.

AI evaluation / LLM capabilities
Diagram showing an AI agent updating its understanding of a complex environment
Jun 29, 2026 · 8 min

Why a smart AI needs a self-evolving map of the world to get things done

Discover how a new AI approach, WorldEvolver, enhances agent reliability by revising its internal world model, leading to up to 37% higher success rates in complex tasks.

AI agents / LLM
A spreadsheet with numbers and an AI symbol looking at it, with some numbers highlighted as errors.
Jun 29, 2026 · 7 min

Why AI sometimes misreads numbers in your spreadsheets

Large language models can make 'data referencing errors' when reading tables, but a new approach improves accuracy by up to 12.0%.

AI / LLMs
Diagram showing an AI agent with an internal immune system defending against threats.
Jun 29, 2026 · 7 min

Why your AI agents need an immune system

Autonomous AI agents can be hijacked at runtime. A new defense, the Agent-Native Immune System, offers internal protection, reducing operational risks by up to 37%.

AI security / AI agents
Diagram showing an AI agent oversharing sensitive data with external tools.
Jun 29, 2026 · 6 min

Why your AI assistant might be sharing too much private data

New research reveals that 9 widely used AI agents often overshare private business data, even when completing tasks successfully, posing a significant risk.

AI agents / data privacy
A human hand pointing out a subtle detail on an image displayed on a tablet screen, symbolizing precise AI evaluation.
Jun 29, 2026 · 7 min

Why your AI's vision might be failing even when it seems smart

New research shows AI models can miss critical visual details, revealing an 8% perception gap between open and proprietary systems, despite high benchmark scores.

AI evaluation / multimodal AI
Diagram showing Cara's AI architecture on AWS
Jun 26, 2026 · 7 min

How insurance brokerages are using AI to get work done faster

Insurance brokerages are saving about 10 hours per user per week with AI that understands their specific business needs, not just general chat.

AI / insurance
Diagram showing an AI agent planning and executing tasks using external tools.
Jun 25, 2026 · 8 min

How AI agents are doing more than just talking

AI agents represent a major shift, moving beyond chatbots to autonomously perform complex tasks, potentially reducing operational time by 37% or more.

AI agents / business automation
Diagram showing Copilot in Excel connecting to skills and financial data sources.
Jun 25, 2026 · 7 min

How AI agents in Excel can automate complex finance workflows

Microsoft's Copilot in Excel uses AI agents and custom skills to automate multi-step financial tasks, connecting to over 6 financial data providers.

AI / Finance
Diagram showing how an AI agent uses tools and reasoning to complete a task.
Jun 25, 2026 · 8 min

How NVIDIA's Nemotron 3 Ultra helps AI agents get more done

NVIDIA's Nemotron 3 Ultra, a new 550B-parameter AI model, helps AI agents complete complex tasks 5x faster and with 30% lower cost.

AI agents / NVIDIA
Flow diagram showing a small AI model guiding a larger AI model's development process.
Jun 24, 2026 · 8 min

How a small AI model can speed up building bigger, smarter ones

A new research framework called Knowledge Cascade can cut the computational cost of developing large AI models by up to 37%, making advanced AI more accessible.

AI strategy / machine learning
Diagram showing an AI agent interacting with a computer screen to perform tasks.
Jun 24, 2026 · 8 min

How AI can now use your computer to get things done

A new AI capability, 'computer use' in Gemini 3.5 Flash, lets AI agents perform tasks across applications, potentially cutting manual work by 37%.

AI agents / Gemini 3.5 Flash
Diagram showing an AI agent handling a healthcare appointment call.
Jun 24, 2026 · 8 min

How an AI agent can handle healthcare appointments by voice

Learn how an AI voice agent, powered by Amazon Nova 2 Sonic, can reduce healthcare appointment no-show rates by 5-30% by automating patient conversations.

AI agents / healthcare AI
Diagram showing Huntington Bank's document redaction workflow using AWS services
Jun 24, 2026 · 7 min

How banks can find and hide sensitive customer data in millions of documents

Huntington Bank used AWS services to automatically find and redact sensitive data from over 400 million documents in just months, not years.

AI / Data Security
A dashboard showing GPU utilization metrics across a Kubernetes cluster.
Jun 24, 2026 · 8 min

Why your AI investments might be sitting idle: understanding GPU usage

Many companies underutilize their expensive AI hardware, with up to 75% of GPUs sitting idle. A new monitoring tool provides real-time visibility.

AI Infrastructure / GPU
How AI agents pay for things on their own
Jun 22, 2026 · 6 min

How AI agents pay for things on their own

AI agents are starting to pay tiny amounts per task instead of running on a flat subscription. A look at how Ampersend and Amazon Bedrock make that work safely.

AI agents / payments
Diagram comparing traditional air cooling to liquid cooling for AI servers.
Jun 22, 2026 · 8 min

Why warmer coolant makes AI data centers more efficient

NVIDIA's new AI servers use 45°C liquid cooling, cutting data center energy use by up to 40% and saving a 50-megawatt facility over $4 million annually.

AI infrastructure / data centers
Diagram showing an AI agent using a web search tool to get up-to-date information.
Jun 19, 2026 · 7 min

Why an AI agent needs web search to stay current

Discover how Web Search on Amazon Bedrock AgentCore updates AI agents with fresh information in minutes, enabling better real-time business decisions.

AI agents / Amazon Bedrock
Amazon Bedrock AgentCore harness: from idea to a working agent
Jun 18, 2026 · 6 min

Amazon Bedrock AgentCore harness: from idea to a working agent

Amazon released a tool that handles the setup work behind AI agents. Here is what an agent is, what the harness does, and why the timing matters.

AI agents / enterprise AI
How ChatGPT Enterprise now shows what your team spends
Jun 18, 2026 · 6 min

How ChatGPT Enterprise now shows what your team spends

OpenAI added usage analytics and spend controls to ChatGPT Enterprise. Admins can now see who uses what and set monthly limits before costs climb.

enterprise AI / ChatGPT
Agentic resource discovery: letting AI agents find their own tools
Jun 17, 2026 · 7 min

Agentic resource discovery: letting AI agents find their own tools

AI agents can usually only use tools a developer set up in advance. A new shared standard lets an agent search for the right tool on its own, mid-task.

AI agents / tool discovery
AI agents that write their own code to answer questions about company data
Jun 17, 2026 · 7 min

AI agents that write their own code to answer questions about company data

A new paper shows AI that writes and fixes its own code to read a company's data, tested on seven SQL benchmarks, matching or beating the best results.

AI agents / enterprise data
AI is helping attackers move faster, and what it means for you
Jun 17, 2026 · 6 min

AI is helping attackers move faster, and what it means for you

The same kind of AI behind ChatGPT now helps online attackers work faster. A Microsoft post explains how, and what a normal person can do to stay ahead.

AI security / cybersecurity
Can AI agents be trusted in drug discovery? What TxBench-PP found
Jun 17, 2026 · 8 min

Can AI agents be trusted in drug discovery? What TxBench-PP found

A new test gave 11 AI models real drug-lab data. The best one got the right answer 59.3 percent of the time, and none of them could be trusted on their own.

AI evaluation / drug discovery
LifeSciBench: how OpenAI measures AI on real science
Jun 17, 2026 · 6 min

LifeSciBench: how OpenAI measures AI on real science

A plain-English look at LifeSciBench, OpenAI's new test that checks how well AI handles real life science research, and what its low scores tell us.

AI evaluation / life sciences
Why AI agents give wrong answers about your company data
Jun 17, 2026 · 6 min

Why AI agents give wrong answers about your company data

An AI agent is only as smart as the context it can reach. AWS Context maps how a company's scattered data connects so agents stop guessing.

AI agents / data
What the Amazon Bedrock InvokeGuardrailChecks API means for running AI agents safely
Jun 16, 2026 · 8 min

What the Amazon Bedrock InvokeGuardrailChecks API means for running AI agents safely

Amazon released a new safety tool on 16 June 2026 that checks an AI for risk at every step of its work, not just at the start, and gives each check a score instead of blocking on its own.

ai-safety / agentic-ai
Can AI models remember what they saw? What RNG-Bench found
Jun 16, 2026 · 7 min

Can AI models remember what they saw? What RNG-Bench found

Researchers built a test to check whether 7 AI models can act on facts that are no longer on the screen. The best one got 62.3 percent of card pairs right, and bigger boards broke them fast.

large language models / AI evaluation
How companies actually get value from AI
Jun 16, 2026 · 6 min

How companies actually get value from AI

A plain guide to what a company has to figure out before AI pays off, built around three questions every business is asking about it right now.

enterprise AI / AI transformation
Why AI apps can be slow to start, and how caching speeds them up
Jun 16, 2026 · 6 min

Why AI apps can be slow to start, and how caching speeds them up

When an AI app gets busy, it has to wake up new computers to keep up. Those computers download a big file first, and that wait is what caching now cuts.

AI infrastructure / cloud computing
CEO-Bench: can AI agents run a company for 500 days?
Jun 15, 2026 · 6 min

CEO-Bench: can AI agents run a company for 500 days?

A Princeton test puts AI agents in charge of a fake software company for 500 days. Most went broke. Two finished rich. Here is what that tells you.

AI agents / benchmarks
Why AI agents fail in production and how failure detection fixes it
Jun 15, 2026 · 7 min

Why AI agents fail in production and how failure detection fixes it

When an AI agent gets something wrong, the mistake usually hides several steps back. A new AWS tool reads the agent's own logs and tells you which step broke and why.

AI agents / observability
Teaching AI to use a computer by letting it practice
Jun 14, 2026 · 6 min

Teaching AI to use a computer by letting it practice

Most AI only chats. A new report shows how teaching an AI to click and type its way through a real desktop made it far better at finishing tasks.

AI agents / computer use
How Rocket Close cut title work with an AI that takes actions
Jun 12, 2026 · 7 min

How Rocket Close cut title work with an AI that takes actions

A real company put an AI agent on its slowest, most paperwork-heavy task and cut contact center calls and emails by 30 percent. Here is what it did, in plain terms.

AI agents / agentic AI
How to tell if an AI agent is actually working
Jun 11, 2026 · 6 min

How to tell if an AI agent is actually working

A right answer can hide a broken AI agent. A new AWS toolkit checks the steps, not just the final reply, and one test scored faithfulness at 32 percent.

AI agents / AI evaluation
Turning piles of documents into clean data without typing it in
Jun 11, 2026 · 6 min

Turning piles of documents into clean data without typing it in

How AI reads scanned PDFs and pulls out the useful facts on its own. A plain walk through fast one-at-a-time versus cheaper overnight processing for beginners.

document AI / AI infrastructure
The price of anarchy: why serving AI cheaply is a traffic problem
Jun 10, 2026 · 6 min

The price of anarchy: why serving AI cheaply is a traffic problem

A new paper studies what happens when an AI service gets crowded. Above a certain load, costs grow more than 280 times. Here is why, in plain words.

AI infrastructure / AI methods