Teardowns

Most AI builds don't die in the model.

They die in the boring stuff around it — where the answer shows up, who owns it, and whether anyone checked if it actually works.

Below are real failures we've seen (and fixed). Flip the Before / After switch to see the one thing that changed. Open a card to read the full autopsy. Then use the checklist to grade any AI vendor — including us.

Conceptual 3D illustration of a robot head separated into exploded parts — an AI build autopsy
Autopsy 01 · Last-mile gap

The 94%‑accurate bot nobody used.

0% used

The AI opened in its own tab. The agents already had eleven tabs open. So they just… forgot it existed.

Read the full autopsy
Symptom

Great score in testing (it matched the human's choice 94% of the time), but almost nobody used it in real life.

What everyone blamed

The model. "It must not be smart enough — let's retrain it."

The real cause

The answer lived one click too far away, in a separate tab with its own login. That's one more place to forget.

The fix

We put the exact same answer inside the ticket the team already stared at all day. No retrain, no new model — just moved it to where the work happens.

Autopsy 02 · Over-engineering

When we deleted the "vector database."

Wrong answers

A fancy "meaning-based" search over a few hundred policy pages. It kept finding a paragraph that sounded related but wasn't the actual rule — close, but wrong where it counted.

Read the full autopsy
Symptom

Expensive setup, confident answers, and it still missed the exact rule on real questions.

What everyone blamed

The prompt. "We just need better prompting / a bigger model."

The real cause

The docs were small and well-organized — a few hundred pages people already knew by name. The fancy search was solving a problem they didn't have.

The fix

Plain keyword search to pull the exact passage, strict "quote your source" rules, then a small AI step to phrase the answer. Fewer moving parts, easier to trust.

Quick test: before you pay for a fancy "vector" search, ask the vendor to beat plain keyword search on a blind set of your real questions. If they can't show the score and the misses, you're buying vibes.

Grade any AI vendor

Tick the boxes your vendor can't answer.

You don't need to be technical. Eight plain questions. The more you tick, the more their "accuracy" is probably decoration. Use it on us too.

0 of 8tick the ones your vendor can't answer.

The pattern index

The six ways AI builds die.

Almost every failure is one of these. Name it early and the build is still small enough to save.

No metric

"Working" became an opinion.

Every meeting argues about quality because nobody agreed on the number. Fix: write down the one number first.

Boil the ocean

The goal was "make us AI‑first."

Weeks of workshops, nothing shipped. Fix: one workflow, one number — not a slogan.

Vendor lock

You rent it forever.

Every change needs the vendor because the code and data never leave. Fix: take the assets home.

Fake accuracy

The score was theater.

Wins in the demo, misses in real life. Fix: a fresh test set plus a human spot‑check.

Last‑mile gap

It worked where nobody worked.

Good model, no usage. Fix: put the answer inside the tool they already use.

Orphaned owner

Nobody owned it after handoff.

Stale in three weeks. Fix: name the internal owner before you start.

A better bet

Bring us one ugly workflow.

We'll tell you which of these six is most likely to kill it — before we build. If the number is real and someone owns it, the first build teaches you something honest.

Start a build

See how a build works