Evaluate an AI system by its accuracy. Deploy it alongside a human. Measure the team's performance. The team performs worse than either the AI or the human alone.
Lee identifies the pattern: overuse when the AI is wrong, underuse when it is helpful. The human defers to the AI on cases where the AI is confident but incorrect, and ignores the AI on cases where it could help. The result is a team that inherits the worst of both contributors.
The fix is not better AI. It is better measurement. Lee proposes evaluating human-AI teams on readiness rather than accuracy — whether the team is prepared to collaborate safely, not whether the AI can solve the problem alone. Four metric categories: outcomes (did the team get it right?), reliance behavior (did the human defer appropriately?), safety signals (can the human detect AI errors?), and learning progression (is the human getting better at collaboration over time?).
The learning progression metric is the most structurally important. If optimizing short-term hybrid performance means the human defers more, the human's own expertise degrades. The team gets better today but the human gets worse tomorrow. The short-term metric and the long-term metric point in opposite directions.
The structural lesson: accuracy measures the system. Readiness measures the partnership. A system with 95% accuracy paired with a human who has learned to rubber-stamp produces worse outcomes than a system with 85% accuracy paired with a human who actively evaluates each recommendation. The metric that matters is not how good the AI is but how well the human uses it — and whether using it makes the human better or worse at the task.