Skip to content
Field notes

What we’ve learned shipping this.

No think-pieces, no predictions. Implementation detail from agents running in production right now — what worked, what broke, and the numbers behind both.

AgentsEngineeringOperationsReliabilitySecurityStrategy
Engineering6 min read

Evals: how to actually know whether your AI system works

Vibes-based testing is why AI projects stall at 'promising demo'. A practical guide to building eval sets, choosing metrics, using LLM-as-judge without fooling yourself, and catching regressions.

Nelson Archer · 19 Jun 2026Read
Engineering6 min read

Cutting LLM costs by 10× without cutting quality

Prompt caching, model routing, batching and context discipline. Four levers that reliably take an order of magnitude off an AI bill — with the mechanics of why caching silently fails.

Nelson Archer · 21 May 2026Read
Operations6 min read

Measuring AI ROI without lying to yourself

Most AI business cases are built on hours saved that never left the payroll. A method for measuring what actually changed — baselines, counterfactuals, and the four categories of real return.

Nelson Archer · 12 Mar 2026Read
Strategy6 min read

Why your AI pilot never shipped

The demo worked. Nine months later nothing is in production. Seven structural reasons AI projects die between staging and live — and what the teams that ship do differently.

Nelson Archer · 11 Feb 2026Read

Want this applied to your operation?

A 30-minute audit call. We map your workflows and name the three highest-leverage agents, with what each is worth.

Book an audit call