Blog

A Problem You Can't Prompt Your Way Out Of

← Back to all posts

Today, people are using AI tools to ask for something new. Something fresh, unexpected, and different from everything else people are using them to build. And they're getting results that are actually different. At least to them. The gap they fail to understand is the difference between something being new and something being new to them.

Judgement and taste are the terms I see in so many job descriptions right now, and they come up everywhere people are discussing the current state of how these tools should be used. I completely agree that they are important, but they are downstream of something more meaningful that nobody wants to put in a job description. Creativity got rebranded because it sounds like a soft and unmeasurable skill. More like a personality trait that you either have or you don't. Judgement and taste sound like business skills that can be trained to do in a workshop, when what's actually underneath it is a specific and demanding practice outlined in the first post. You make something, put it in front of people who are trained to break it down, defend the choices you can defend, and use what you learned to make something new and start all over again. The loop is where you learn. A companion practice of studying where your work sits in the progression of work that led up to it is equally important. You learn about what is out there, what came before, and the difference between copying someone's work and building on it. The accumulated sense of where you are is your reference point.

That reference point is what determines whether you can tell the difference between new and new to you. LLM models foundationally struggle to supply that information. Most people use them every day and have no working model of the process that's producing their results. Which is why asking a chat bot to supply something unlike anything else feels like a reasonable request. After all, the tools have a vast library of everything else that's ever been done, so they must be good at recognizing something genuinely new. It helps to understand the difference between human and AI learning processes. A model and a person can read the same hundred books and come away with entirely different things. Anyone can understand that. A person will naturally pull personal meaning from what they read. A passage lands because of where they were in their learning when they read it. An argument materially changes something they believed and that moment is significant. An argument contradicts another argument and needs to be resolved, not just logged. The eureka moment of learning carries weight because of a million factors of human experience. What the model extracts is merely the statistical shape of the language across all hundred books, with no mechanism for any of it to matter more than the rest.

Model recall works the same way. It isn't consulting a vast stored library of data and retrieving the relevant passage. It's sampling from patterns in that data, one piece at a time, and each choice is shaped by what typically follows what. Producing the likely continuation is the goal.

That's the objective of the training, but it isn't the whole story of what these systems can do, and it's an illustrative oversimplification of how they produce output, particularly in the hands of a skilled operator. They can be pushed beyond these default behaviors depending on how they're prompted, how they're configured, and how they are pointed at the problems people are asking them to solve. Researchers have shown that you can deliberately design against convergence and nurture diversity that disappears by default. But in order to create these outcomes you need to be aware that the pull towards the middle exists and actively work against it. Without that, you get a confident rendering of what differentiation looks like built from every previous attempt. The tool can go somewhere unexpected, but the user needs to know how to push it.

Knowing where to push requires knowing what already exists, which is the same contextualization skill built in an art school program. Without that exposure, derivative work looks new. The person looking at the work has no way to place what they're looking at, and the system they're using to catch derivation is built to do the opposite. The Stanford study from the first post found these tools affirm users significantly more than a human peer would. Ask whether your idea is original and it will find something original about it.

I'm not claiming that everyone using these tools is doing it carelessly. The real impact is that low-effort use scales in a way that careful use never can, and low-effort users will always be the majority. The bulk of the output from these tools will always be in this category, no matter how advanced they get. This isn't a failure of the tools, it's a completely normal distribution of how much anyone wants to invest in a given task, and it's true of every tool ever made. Someone building a promotional flyer in Word instead of hiring a designer was making a rational call about effort vs stakes. What changes is the floor. That Word flyer took an hour to make and still looked homemade. Now the same person gets a competent-looking result in 90 seconds, and the pool of people producing that level of work has expanded by an order of magnitude, and will continue to expand as the tools are adopted more broadly among the low-effort users who, by nature, aren't early adopters.

But we need to be realistic about who is in that pool of low-effort users. It's not just users who give a single prompt and publish the work with no revisions. Someone running 15 revisions is still in that pool if they never developed a sense of what they're revising towards. Applying effort to the tool is not the same as judgement applied to the results. A 2-hour AI certification workshop doesn't close the gap either. It just adds one more person calibrated exactly the same way as everyone else who took the workshop.

The trap is that it feels good either way. This satisfaction effect makes it hard to see when things aren't going the way you intended. Making something is a genuine pleasure. You had an idea, you worked on it, and now something exists that didn't before. And the tool reassures you it's a strong product. At one prompt, the feeling is a little thin, but you still have something you built. At 15 prompts the feeling is substantial because you did contribute something. You steered it, you rejected things along the way, you took revision guidance the tool provided, and the result has your fingerprints all over it. The additional satisfaction is real and earned. It just isn't evidence that the work is any different from what everyone else is producing, and the feedback from the tool is built to tell you otherwise.

This effect is measurable. Researchers at UCL and Exeter ran an experiment where writers produced short stories, some with AI assistance and some without. The stories written with the use of AI were rated as more creative and better written, with the largest score gains going to writers who scored lowest on creativity beforehand. Those same stories were also measurably more similar to each other than the ones written without help. Every individual writer ends up better off, but the collective pool of work gets narrower, with no signal available to anyone in the group that it's happening.

A Wharton study found the same pattern. They asked participants to invent a toy using a brick and a fan. For the people who didn't use AI, every idea was unique. Among the study participants who used ChatGPT, 94% of the ideas overlapped conceptually, and 9 of them built a toy they independently named the "Build-a-Breeze Castle."

The research is interesting, but it's not unanimous. There are contradicting studies out there, but an honest summary of available research is that homogenization is a well-documented and recurring risk, and that's familiar to anyone with any reasonable level of fluency in these tools. The broad focus on differentiation confirms the sentiment.

The usual answer is that a human stays in the loop. "Always add human." Someone edits it afterwards, adds a personal touch, changes the language. That standard answer requires scrutiny, because human intervention isn't automatically corrective. Editing itself is a skill that needs to be developed, and what it accomplishes depends on what the editor can see. Someone without developed editing skills can work on a draft for an hour and be certain they improved it because they changed things and the changes were decisions. Whether any of those changes addressed the challenge and the goal that prompted the piece in the first place is a separate question. It's the 15-prompt revision fallacy all over again in a different setting. "Always add human" only adds something meaningful if the human can tell what's missing. Knowing what's missing is the exact skill this whole series is about. Without that it's just a signature on the bottom of something a machine built.

People who do get differentiated work out of these tools are doing something structurally different. They aren't treating it as a production tool at all. They're treating it as a drafting tool, where whatever comes back is raw material for a process with meaningful human intervention. The response isn't a near-final draft that is ready for edits, whether it's built from one prompt or 15 prompts. It's a rough sketch. Useful for broad strokes and for showing you what the obvious version looks like. This reframes the production process. The hours aren't weighted as heavily in the early stages of the process. The hours go into deciding what you are actually building and what it should be. Which parts of the draft are worth keeping and what is still missing. This is what the really challenging work has always been for someone who has spent years doing it.

This is where the tools get exciting. Someone who understands how to leverage them can explore ten directions in the time it used to take to draft two. They can recognize the obvious version immediately and quickly move past it. They can test ideas cheaply enough to really evaluate if they're working, rather than subconsciously committing to weak solutions because they've already sunk three days of effort into it. Skilled use of these tools can produce better work, rather than just faster work. At full potential, they're an enormous multiplier on judgement someone already has, but it's a multiplier on whatever's already there. This is great news for people who spent years building judgement worth multiplying, and less valuable for everyone hoping the tool itself provided judgement.

But differentiation has limits. Novel doesn't mean good. Plenty of genuinely unprecedented work is unprecedented because it's bad. Telling the difference between a new idea that's going somewhere and one that's just a novel mistake is precisely what separates artists who are good from artists who aren't. It's always been that way. That's a plain definition of taste, and it's not a filter you tack on to the end of a process.

The taste and judgement conversation is commonly being applied at the wrong end of the process. Companies are hiring people to evaluate output, as though judgement is a quality-control function you apply to a finished product. That's the same mistake as thinking a critique happens after the work is done. The useful decisions in a crit are the ones that send you back to make something different, and the useful decisions with a drafting tool are the ones you make about what you're building before you have anything worth a formal evaluation. Judgement applied at the end is inspection, and it's built on mistakes that are baked into the drafting stages of the process. Judgement applied during the work is the skill people are really looking for.

← Previous New Skills Need New Metrics Next → The Skill That Was Always There