KENNY CHIEN / IDEAS / SLOP IS A CHOICE
Slop is a choice
AI does not produce mediocrity. Unexamined taste does.
Generation is free now. Judgment is not. Slop is what happens when you confuse the two.
In one paragraph
Slop is a choice means that AI-generated mediocrity is a property of the process around the model, not of the model itself. The same model yields slop on one team and production-grade work on another; the variable is whether anyone with taste examined the output before it shipped. Slop appears when generation is treated as completion — no review, no explicit quality bar, no evals. The antidote is judgment applied systematically: taste to know what good looks like, review culture to enforce it on every artifact, and evals to make that enforcement repeatable at machine speed.
The argument
Where slop actually comes from
A language model generates the median of everything it has seen. That is its nature and, correctly used, its virtue: a competent first draft of almost anything, in seconds. Median is the default. It was never the destiny. Blaming the model for slop is blaming the printing press for bad novels.
Before AI, mediocre work had a natural rate limiter: it still took effort to produce, so there was less of it, and the effort forced at least a minimum of examination. Generation removed the limiter. Volume with judgment is leverage. Volume without judgment is slop — and the volume was chosen by whoever pressed generate and shipped the result unread.
Taste is the scarce input
Taste is unglamorously concrete: knowing what good looks like in your domain, and being able to say why. It is not a gift. It is trained — by reading excellent code and excellent prose, by being reviewed hard, by rewriting until the difference between fine and right becomes visible.
The economics have inverted around it. Generation is now nearly free; discrimination is not. The scarce skill is the eye that rejects nine drafts and can articulate why the tenth is right. A person with taste plus a model produces more good work than either alone. A person without taste plus a model produces confident slop at unprecedented speed. The model is an amplifier, and amplifiers do not care what they amplify.
Review and evals are the antidote
Taste in one head does not protect an organization. It has to be institutionalized, and there are exactly two mechanisms that work.
The first is review culture: human judgment applied to every artifact that ships, with AI-assisted work held to the same bar as human work. The moment a team maintains a separate, lower standard for generated output, it has chosen slop and outsourced the signature. The second is evals: taste written down as executable criteria — golden examples, failure taxonomies, acceptance thresholds — run against every generation and every model upgrade. Review scales judgment across people. Evals scale it across volume. Teams that build both ship AI-assisted work indistinguishable from their best human work, because it was judged by the same standard.
What it means for your team
Set one quality bar. If AI-assisted work gets a softer review, you have chosen slop as policy.
Make review non-optional and named: a human signs everything that ships, and the signature means they read it.
Write your taste down as evals — golden examples and failure taxonomies your pipeline runs automatically.
Train taste deliberately. Put good and mediocre outputs side by side and make the team articulate the difference.
The steelman
The strongest counterargument is arithmetic: if models let a team generate a hundred times more, humans cannot review it all, so slop wins by volume no matter what culture you build. Correct — if review means reading everything. It does not.
Manufacturing solved this a century ago: quality is controlled with sampling, tight specifications, and instruments — not by inspecting every unit by hand. The equivalent here is reviewing at the shipping boundary rather than the generation boundary, sampling deeply, and pushing the rest of your judgment into evals that run at machine speed. And the arithmetic hides a choice: nothing obliges you to ship everything the model produces. Teams drowning in their own AI output chose the volume before they chose the quality bar. Choose in the other order.
Where to go next
Worried your team is shipping slop?
Tell me how AI-assisted work gets reviewed on your team today. I will reply with the first hole to close.