The Field Notes Library

What good oversight looks like when it breaks.

Six failure modes I've seen across every AI deployment I've audited. Each one is a specific way human oversight silently degrades in production. Tap through them, then flip a card for what to look for in your own system.

Failure mode · 01/06High risk

Warm Body Problem

A human sits in the review loop. Nobody checks whether they are actually looking. Compliance documentation says oversight exists. The error rate says otherwise.

Review queuetap Approve
Item 1 of 24
Observed across 40+ deployments
Failure mode · 02/06Managed

Burnout by Design

Review volume grows with the AI's output. Human capacity does not. Quality degrades quietly over a shift, then over a quarter.

Reviewer loaddrag the slider
Items per reviewer per day: 20
Review quality

Sustainable

Observed across 40+ deployments
Failure mode · 03/06High risk

Gate Fatigue

Add checkpoints everywhere and humans stop reading them. Too many gates produces the same failure as no gates: automatic approval.

Approval gateskeep tapping
Gate 1: Approve this action?
Observed across 40+ deployments
Failure mode · 04/06Managed

Trust Accumulation Risk

Familiarity raises approval rates regardless of actual reliability. Trust grows on autopilot while accuracy stays flat.

Sessions togethertap +
50 sessions
Auto-approve rate
22%
Actual accuracy
78%
Observed across 40+ deployments
Failure mode · 05/06Critical

Silent Failure Accumulation

Errors nobody catches do not disappear. They accumulate in logs nobody reads, until one surfaces as an incident.

System statuscheck the log
All systems normal
Observed across 40+ deployments
Failure mode · 06/06Critical

The Measurement Gap

You cannot govern what you cannot measure. Most organizations measure AI capability constantly. They measure nothing about oversight quality.

Two meterstap Instrument
Model performance
94.2%
Oversight quality
no data
Observed across 40+ deployments