Ask a manager how the team is performing and they'll pull up a dashboard. Ask someone if their plan is good, and increasingly they'll paste it into a chatbot. Both hand off the same questions to something confident, fast, and never actually confirm whether it answered the question.
Two things about this moment are already obvious to anyone paying attention to the current conversation. Judgement and taste have become the skills everyone claims to want. In February 2026, Paul Graham posted that in the AI age, taste has become more important. When anyone can make anything, the differentiator is what you choose to make, linking back to an essay he'd written in 2002. Two days later, OpenAI's president Greg Brockman added five words: taste is a new core skill. Cloudflare's CTO had already written in January that taste is the engineering differentiator for 2026. Within a week Business Insider was covering the debate, and by March the New Yorker was asking why tech was now so obsessed with the word. I think what was driving that was people finding a flattering name for what they already believed about themselves.
It also quickly showed up in job postings. PwC reviewed more than a billion of them this year and found entry-level roles most touched by AI are seven times more likely than untouched ones to ask for judgement and leadership skills that PwC describes as traditionally linked with postings for senior roles. The same report also points out that the junior work, where people acquired that judgment through practice and partnering with someone more experienced, is quietly getting automated away. What nobody has proposed is a plan for what replaces it.
People have been quick to notice and discuss the problems with this, but I think many people are missing the deeper systemic issues. Writing judgement and taste into a job description takes one ten-minute edit. Building an actual way to check for it, in both the hiring process and a performance evaluation, is an entirely different problem, which lands on top of a system of performance metrics that is already being challenged in the AI era.
There's a deeper problem with performance metrics, and it has nothing to do with AI. A campaign that performs well gets read as proof the creative behind it was good. But performance is never caused by one thing. It's the value of the product, the strategy behind the campaign, the timing of a social post, who got targeted and whether they were already primed to want it, a trend nobody planned for, a competitor stumbling that same week. The work itself, the writing, the design, the execution, is one ingredient in that mix, and a performance number can't isolate how much of the outcome belongs to it. Nielsen's analysis of nearly 500 ad campaigns found creative responsible for 47 percent of the variation in sales, the single largest factor, ahead of reach at 22 percent, brand at 15 percent, and targeting at 9 percent. Even the biggest lever in the system explains less than half of what happened. A number saying a campaign performed well is not the same as a number saying the work was good. It's a confluence of causes a dashboard was never built to separate, and AI didn't create that problem, it's just raising the stakes on organizations that lean on outcome data even harder while judgment itself resists measurement
A research group called METR ran a careful study on whether AI actually speeds developers up. What they found was that it slowed them down by 19%, while the developers themselves claimed they were 20% faster. A year later, METR tried to run the same study again using the same method and the same design. They couldn't. Too many developers refused to take part unless they were allowed to use AI for the whole study. The pool of people willing to be measured the old way had already changed. The method failed, but the issue wasn't sloppy research. The situation changed enough in 12 months that it broke the comparison the whole study depended on. The application of the roles is shifting faster than the metrics meant to measure them.
Faros, a firm that tracks engineering data across thousands of teams, found that companies with heavy AI use produced code that is bigger and sloppier. Mistakes are slipping through more often, and the time spent checking the work vs making it is now five times longer than it used to be. When making gets cheap, checking becomes the job. And the metrics haven't caught up. Most teams are still measuring how much got made, while the data that actually matters remains a mystery.
Updating that apparatus is not a trivial task. Research from The NeuroLeadership Institute found that 88% of companies that overhauled the way they evaluate people took two years before the new approach gained significant traction. The job posting changes in ten minutes. The system behind it runs on a multi-year timeline, and anyone who's been through a performance review systems change knows that rings true.
When organizations identify that gap they reach for ways to close it, and can mistakenly grab for something that's easy to work into a dashboard. This year, the headline blunder was tokenmaxxing. A Meta employee built an internal leaderboard called Claudeonomics that ranked roughly 85,000 colleagues by how many AI tokens they used, handing out titles like "Token Legend" to the top user, who burned through 281 billion tokens in a single month. Fortune estimated that at the cheapest available rate, that one employee could have cost Meta more than $1.4 million. The leaderboard came down two days after The Information reported on it, and Meta has since eased off token counts in performance reviews. Counting the activity felt like a metric for success, while encouraging bad behavior leading to worse business outcomes.
The traditional way of measuring judgement and taste is now getting squeezed from both ends. Companies used to rely on borrowed judgement. You don't need to have taste and judgement yourself if you can identify someone who has it and defer to them. Junior hires weren't expected to onboard with taste and judgement. They were expected to pass decisions through someone whose taste and judgement had already been tested and to pick it up along the way. The routine work that taught people what good looks like is being automated. What's being cut isn't a perk, it's the only mechanism we had for transferring this particular skill, that's now being listed as a requirement in job descriptions.
The obvious replacement is to trust the experienced people who are left. But, when judgement can't be measured, evaluation relies on an individual's personal read, and a systematic review of workplace studies found that people reliably favor hiring people who resemble themselves unless something actively works against it. Nobody feels that bias, it feels like recognizing quality. Performed on a series of hiring and evaluation cycles and an organization converges on a single kind of person. This is the same outcome researchers documented in algorithmic screening tools. Employers sharing one system produced the same rejections everywhere without a single decision seeming unfair. It sped up the process, but didn't produce different results.
This doesn't mean the old metrics or the skills built around them stopped mattering. A track record of good output still tells you something, you just have to determine what it's actually telling you. Judgment doesn't replace measurement any more than an art degree from the last post replaces a business degree. But the systems designed to recognize this skill are going to take time to catch up, and it's going to be longer than anyone updating a job description would like. Two years implementing an organizational change in evaluation is the fast case, not the slow one. Confidence in the current process is exactly what makes the problem hard to notice.
If you're looking to explore solutions, look again at what the Faros data actually describes. The production got cheap and checking became the job. That means the hours that used to be allocated to production are already free. Right now, most of that freed time is pushed straight back into production, because production is the metric that dashboards measure. It doesn't have to go there. That's valuable time to rebuild the lost mentorship that builds junior hires into senior contributors, but not if they're being evaluated on outdated metrics.