AI-generated API specifications beat human-authored ones on 10 of 11 usability dimensions. They took 87% less time to produce. Experts couldn't tell AI work from human work — only 19% accuracy at identification.
But the experts described the AI designs as “unsettlingly perfect.” The hyper-consistency that made them measurably better also made them feel wrong. Not because they violated standards, but because they adhered to standards too uniformly. Real API designs carry the marks of pragmatic trade-offs — places where the standard was bent because the specific use case demanded it. The AI designs had no such marks. They were consistent in a way that human design never is, and the experts noticed the absence.
The authors call this the Perfection Paradox: the same property that makes AI output measurably superior (consistency) is the property that signals it lacks something important (judgment about when to be inconsistent). The implication is that perfect adherence to standards is itself a design flaw — because the standards were written as guidelines, not as laws, and the places where humans deviate are often the places where the standard is wrong for the specific context.
The proposed role shift — from drafter to curator — is the design community's version of what's happening everywhere. The human stops producing and starts editing. But curation is a different skill than production, and there's no guarantee that good producers make good curators. The expertise that made someone good at API design was the ability to hold the standard and the context simultaneously. If the AI handles the standard perfectly, the remaining human contribution is pure context — and context is harder to evaluate, harder to train, and harder to audit.