From the field
Field notes from live engagements.
Methodology, AI assurance research, and reports from inside live engagements. We write the way we work, evidence first, claims second, nothing the data won't carry.
Latest writing
The whole register.
Showing 3 of 3 articles
01
MethodologyEssay
The test pyramid was built for code. Agents need a different shape.
Coverage and pass-rate stop meaning much once behaviour is probabilistic. A field-tested model for what to measure instead.
02
AI AssuranceField Report
What "safe to ship" actually means for an agent in production.
A purpose-bound evaluation isn't a benchmark score. The evidence trail behind one banking assistant, 4,100 adversarial turns and the transcript that held up the release.
03
AI AssuranceNewsletter
When the benchmark says ship and the evidence says hold.
A field note on the one-in-eight problem, a methodology habit we've abandoned, and three things worth your attention this fortnight.
No articles in this topic yet, check back soon.
The newsletter
QE notes from the field.
One considered email every other week, a field report, a method we changed our minds about, and the AI assurance debate the category is avoiding. No digests, no roundups, no fluff.
6,200+QE & eng readers
Bi-weeklyThursdays, 8am
0%Sales emails
Read a recent issue →
Subscribe
You're in. The next issue lands Thursday.