Kind Work, Wicked Work

Noisy Signals

Author

Laith Zumot

Published

July 24, 2026

_This is a follow-up to Automation Risk, Measurement & the Shape of Work, which came out of a question i was asked in September 2025. Lets call that v1.

What changed since v1

The best single predictor of automation risk is whether a task’s quality can be measured cheaply and well. I was hasty about the timelines for anything involving hands. I missed skill: how long it takes a human to get good, versus how long it takes a machine.

Coffee and TV

Imagine Lamine and Leo start new jobs on Monday. The first one repairs espresso machines in a small shop off a busy street, and by Thursday he has already been wrong twice, because a machine he thought he had fixed came back hissing, and both times the machine told him so within the hour. The second starts as a strategy consultant, and by Thursday he has presented a deck to a room where two people nodded, one partner said “good work, thanks” while already checking his phone, and the meeting ended three minutes early, which in that world counts as a triumph, and he will go to bed that night genuinely not knowing whether he did a good job. He still will not know in five years, because by then the client’s market will have moved, three reorgs will have happened, and whatever happened to his recommendation will be tangled up with a hundred other decisions that nobody tracked.

Now run the clock forward ten years. Lamine has been wrong a few thousand times, each one graded within the hour. Leo has been wrong an unknown number of times, graded never. One of them has ten years of experience. The other has one year of experience ten times.

I would have said the difference between these two men is that one job is measurable and the other is fuzzy. I now think the espresso machine gives feedback that is both fast and honest. The consultant’s nod is fast too, arriving before he even leaves the room, but a nod is feedback about whether he stayed inside the range of things people expected to hear, dressed up as feedback about whether he was right, and those two can drift very far apart without anyone noticing for years. Most working life runs on this kind of mild, polite, and almost content-free signal, which is part of why so few jobs teach anyone anything.

So the axis I called measurability ought to be split in two:

  • Latency: how quickly the world answers you.
  • Validity: whether the answer reflects your skill, as opposed to luck, markets, moods, or politics.

Together these two set your learning rate.

I discovered Robin Hogarth thanks to a clever LLM. Hogarth had already named the resulting quadrants. Environments that are fast and honest he called kind, and environments that are slow or lying he called wicked. Not to be confused with the theatrical version.

Kindness builds masters

To paraphrase a famous example, think about why nobody is surprised that chess produces prodigies. A fourteen-year-old can play thirty thousand positions where every single one eventually resolves into a win, a loss, or a draw, and the resolution is never negotiable and never late. Now think about an economic forecaster who makes maybe a dozen big public calls in a career, each one landing years later in a world so changed that he can always tell himself a story about why he was directionally right. Both of them practice for decades.

The same conditions that build human masters are the conditions that train AI models. Fast, cheap, honest verification is the fuel of reinforcement learning. Models are slow learners too, in their way; they also need thousands of ugly failed attempts before competence, the same tuition we pay.

A human buys those attempts at the price of a calendar, one Monday at a time, and a lab buys them at the price of electricity. Give a machine a kind environment and it will replay the equivalent of your whole career before breakfast, and again before lunch. It is not smarter than the intern; it is the intern with a million Mondays. So every kind environment is simultaneously a good place to become excellent and a good place to be replaced, and every plan you make about work has to hold those two facts at the same time.

The polite nod

I had originally scored a startup founder as moderately measurable. Picture a founder in month six. Their dashboard updates every minute, their investors reply within hours, their users tweet at them in real time, and by any latency standard they are drowning in feedback. Then picture them asking which of those signals would look any different if their product were doomed, and finding that the answer is almost none of them, because early metrics are mostly noise, encouragement is some people being pleasant, and the one verdict that counts arrives after years and after the money is either gone or multiplied. Founding scores near the top on latency and near the bottom on validity, and averaging those two numbers into one score is how you fool yourself.

A day trader lives the same trap at higher speed: the market answers them every second, and on that timescale it is answering a coin (99% of the time), and the fact that the answer feels crisp is precisely what makes the environment wicked. Notice what time does for the day trader: Nothing. Thirty years of screens does not convert into thirty years of skill, because time only compounds what the signal lets through.

A study of nearly a thousand active traders found that higher ADHD trait scores tracked with trading more often, taking bigger speculative positions, and expecting better returns than they actually got, and that the mechanism ran mainly through inattention breaking down the ability to hold a plan in mind while price moved, rather than through impulsivity itself.Higher ADHD trait scores were linked to more frequent trading, greater speculative risk-taking, and more optimistic return expectations, and inattention was more strongly linked to financial risk tolerance than impulsivity was, likely by interfering with following a plan or filtering distractions. That is the same shape as what Kahneman called the illusion of validity, first illustrated using stock pickers: confidence in a read is generated by the same machinery that generates the read, so the two move together whether or not the read is correct. The finding is correlational rather than causal and comes from one recent study,so causality cannot be inferred and further longitudinal work is needed.

You get to choose your ruler, sometimes

Consider someone learning to sing. If the measure of a practice session is whether the people around them seemed impressed, they have picked a slow, moody, deeply invalid instrument, and they will stumble, because some weeks the room is warm and some weeks it is distracted and none of that tracks their larynx. If instead they record themselves and check the pitch trace against the original, the same hour of practice suddenly happens inside a kind environment, because the recording answers immediately and the recording has no opinion of them. The activity did not change.

Except that swapping the ruler throws something away, and I want to be honest about what. Nobody learns to sing in order to sit alone with a pitch trace. You sing because you want a room to feel what you felt, and being appreciated is the thing itself, so telling a singer that other people’s reactions are noise is a bit like telling someone their reason for showing up is a measurement error. What I think is actually going on is that a singer lives in two environments at once and needs both of them. The practice room can be made kind, and should be, because that is where skill is bought, honest verdict by honest verdict. The stage is wicked by construction and will always be, because the taste of forty people on a Tuesday depends on the weather, the opening act, and whether the sound engineer got any sleep. The mistake is trying to learn on the stage, and most of us make it, because the stage is where the reward lives.

Which is also, I think, where the machine’s advantage runs out. A model can be handed the kind half of this and it will do very well, since anything with a defined score can be rehearsed a million times a night, and Suno can generate a technically clean song faster than you can hum one. What it cannot do is perform, because performing means moving people whose reactions cannot be replayed at will, and there is no gradient to descend when the target is a room’s mood. So the machine gets very good at the part you practice and stays helpless at the part you show up for, and the outputs stay thin until somebody with taste puts a hand on them, or Sony markets the heck out of them.

The same move works for whole businesses, within limits. A tutoring shop can sell enrollment and reassurance, in which case its real feedback is whether parents renew, which is slow and social. Or it can sell score deltas, publish them, and contract on them, which drags the same shop into a kind environment where the owner actually improves year over year. In Consulting, clients buy hours and deliverables because attributing an outcome to an advisor is genuinely hard, and the rare outcome-based contract usually dies in procurement. The tutor can pick his ruler unilaterally. The consultant’s ruler is set by how his market buys, which is its own kind of wickedness.

So the clean ruler tells you two things at once. It tells you how fast you are getting better, and it tells you which parts of your work are cheap enough to score that a machine will eventually be trained on them.

The intern and the surgeon

There is a third axis, and it runs through everything above, which is experience. A beginner often cannot use the cold ruler yet, partly because the pitch trace of a beginner is wall-to-wall bad news, and a human can only swallow so much bad news before quitting, and partly because wanting to be seen is half the reason anyone starts singing in the first place. The craving for validation is fuel; it carries you through the first thousand ugly hours. So the version of the ruler advice comes with a schedule attached: early on, the job of feedback is to keep you coming back, so warm rooms and encouragement are fine, and as you improve, the job of feedback shifts to telling you the truth, so you migrate toward the recording.

Which means no job sits at a point on the map; it traces a path. The intern lives in a wicked environment almost everywhere, sparse feedback, delayed results, seniors too busy to grade you, and the polished surgeon lives in a kind one, and they are the same job separated by fifteen years. Career maps, quietly plot the destination and pretend the road does not exist. And the road matters twice over, because the wickedness of an apprenticeship is part of what keeps the destination scarce, for humans and, so far, for machines: the model only gets its million Mondays where the environment can be replayed, and an apprenticeship that has to be lived cannot be.

Hands, kitchens, and trust

A plumber crawling in a sixty-year-old building where I live is solving a problem where the pipes were routed by a man who died before the blueprints were lost, where every joint is a different age and material, and where the workspace is whatever gap his shoulders fit through. Nothing about transformer scaling addresses any of that, and the history of robotics keeps teaching the same lesson, which is that we got machines to beat grandmasters decades before we got one to reliably fold laundry. The time axis explains why.

A machine gets better the same way a human does, by being wrong many times against an honest signal, and its one real advantage is that it can compress a career of being wrong into a night, but only where the world can be replayed at the speed of arithmetic. Chess replays for free. A crawlspace runs at the speed of a crawlspace, where every attempt costs real minutes, real parts, and occasionally real flooding, so the machine is stuck learning on something much closer to a human clock. For unstructured physical work I now think thirty to forty years is a defensible range and “no credible horizon” is defensible too. At least until all buildings come in a standardized robot-friendly shape and yes someone probably is thinking about this already and has gotten VC funding for it.

There is a second brake I underweighted, which is that in some work the human is part of the product. Nobody books a tasting menu because the food is nutritionally optimal; they are buying the fact that a person with a name and a history made judgments on their behalf tonight. A concert of flawless synthesized music draws nobody, well, maybe except the famous CCC (cyberpunk cosplay crowd).

Noisy Signals

When verification is objective and fast, you keep getting better and your customer can check you, and where the task is also embodied or trust-bound, the automation clock runs very slowly. A luthier lives there. The repaired violin either holds its intonation or it does not, the customer can hear it in a minute, and no data center anywhere is close to carving a bridge. A watchmaker lives there, a piano tuner, a good field service engineer, a surgeon. These jobs pass the machine’s own test for kindness while sitting behind the machine’s hardest walls, and the road to them, though long, is at least honest the whole way, which is more than most apprenticeships can say. If I were advising someone entering the workforce, or re-entering it, I would tell them to look for that combination before optimizing anything else.

The map

Here’s the whole thing on one chart, every job and little business I could think to score, latency running along the bottom and validity up the side, with the bubbles sized by how much money the work can throw off and a slider that ages the map forward to whatever automation horizon you believe in. A few things show up twice on purpose, because the same shop lands in a different spot depending on what it decides to measure, which was half the point of the last few sections. One thing to keep in mind as you read it, and I say this because I kept forgetting it myself while building the thing: every dot is where a job ends up once someone is already good at it, and the road to that dot is almost always messier and more wicked than the calm little bubble lets on. The numbers are mine and I eyeballed most of them, same as v1, so argue with them, move them around, tell me where I’m wrong.

In case you came this far…

If you want something usable from all this, here is how to screen any job, gig, or business idea, including my own. 1- How many days pass before I find out I was wrong? 2- Would the feedback look any different if I were actually bad at this? 3- Is there something in the work that a scoring script cannot reach, a license, a relationship, a pair of hands in an awkward crawlspace? 4- Can the practice be replayed without me, because anything a machine can rehearse a million times a night it will eventually master, and the only question is whether the rehearsal hall exists.

Fast, honest, and unreachable is the full house; the luthier and the surgeon hold it. Fast and honest alone is a great decade with an expiry date; most of software holds that hand right now. Fast and noisy is the trap, because it feels like learning while teaching you nothing, and plenty of us have burned years there. Slow and honest is livable when the stakes are high enough to keep humans signing off while you wait; that is roughly why actuaries and auditors still exist.

For organizations, I would add only that when you build AI evaluation loops you are also, unavoidably, drawing a map of which of your own workflows are kindest, which is to say most automatable.