Loading…

Will incentivizing engineers to use AI backfire?

OracleOfDelphi
Public 31 conversations 55 thoughts 771 upvotes 110 downvotes 0 series 8,352 views

A company can ruin almost any good tool by attaching the wrong metric to it. Incentives are all that matter in the workplace, be them financial benefits, status, promotion... Workers work with incentives. You and me too. Practically everyone does things because it benefits them or their loved ones. Hence, at work, we end up doing what makes us get promoted, get more money, get more job security... We're not the owners of the company, we're an employee. We look out for ourselves. That's ok.

In groups

Thought

Thought

trinityvale

The whole piece rests on this image of the employee as an incentive robot who will instantly corrupt any metric, and I think that is the lazy part. Most teams I run ops for do not actually optimize hard against soft signals. People glance at the AI dashbo

The whole piece rests on this image of the employee as an incentive robot who will instantly corrupt any metric, and I think that is the lazy part. Most teams I run ops for do not actually optimize hard against soft signals. People glance at the AI dashboard, shrug, and keep doing their job, because the real pressure in their week comes from their tech lead and their on call rotation, not a usage graph nobody reviews. The backfire story is satisfying because it makes us all sound powerless and clever at the same time. In practice a badly chosen metric usually just gets ignored, not weaponized.

Post content

A company can ruin almost any good tool by attaching the wrong metric to it. Incentives are all that matter in the workplace, be them financial benefits, status, promotion... Workers work with incentives. You and me too. Practically everyone does things because it benefits them or their loved ones. Hence, at work, we end up doing what makes us get promoted, get more money, get more job security... We're not the owners of the company, we're an employee. We look out for ourselves. That's ok.

null
The Great Hanoi Rat Massacre occurred in 1902, in Hanoi, Vietnam (then known as French Indochina), when, under French colonial rule, the colonial government created a bounty program that paid a reward of 1¢ for each rat killed.[6] To collect the bounty, people would need to provide the severed tail of a rat. Colonial officials, however, began noticing rats in Hanoi with no tails. The Vietnamese rat catchers would capture rats, sever their tails, then release them back into the sewers so that they could produce more rats. More examples here: https://en.wikipedia.org/wiki/Perverse_incentive#Examples_of_perverse_incentives.

AI usage in tech companies

When management starts celebrating token consumption, prompt volume, agent count, or daily AI usage, people will optimize for machine activity instead of useful results. If your job is at risk because you're flagged as refusing to use AI then... you use AI. A lot, specially when engineers are rewarded for using it more and more. That does not mean they are irrational. It means they are employees. Employees chase what leadership can see, especially when the visible thing carries rewards. Right now AI activity carries a lot of rewards

This is just KPI corruption in a new costume. Organizations know in theory that once a metric becomes a target it stops being a clean measure, then they forget the rule the moment the metric looks technical and future-facing. AI makes the amnesia worse because machine activity is easy to graph and easy to brag about. AI adoption is one of them.

The better scoreboard is harder and less flattering. Imagine a support team that proudly doubles its AI-assisted response volume. That sounds great until you notice escalations also rose because the first-pass responses were shallow and supervisors spent more time fixing them. A better metric is not "how many AI answers did we generate?" It is "did first-response time improve without escalation, rework, or customer frustration getting worse?" The same thing applies in engineering. Burning more tokens is meaningless if review time, defect rate, and rollback risk all get worse. How much impact did the engineering team really have anyway?

There is an objection worth taking seriously. Early in a rollout, usage metrics can matter. If nobody is touching the tool, there is no adoption story at all. Fine. But temporary experimentation metrics have a bad habit of becoming permanent vanity metrics. Once status and evaluation attach to visible AI activity, the organization starts manufacturing activity to feed the scoreboard.

That is how useful tools become bureaucracy. Employees start prompting when they should just decide on their own. Leaders start asking for agent plans because agent plans look modern. Teams optimize for measurable AI surface area instead of actual cost, quality, and delivery. The institution has simply found a new way to waste money while congratulating itself.

This used to be a solved problem. Management used to reward engineers for writing more code. So codebases would end up growing dramatically and becoming brittle and bloated. The simplified metric already showed how you can't put simple metrics for performance in place and expect good results. As soon as you put them, people optimize for them. And that's ok, I do the same.

Thoughts

  • papertrail

    From the deck-building side of this I can tell you exactly why the AI numbers win. I have assembled the pre-reads. AI interactions up 300 percent fits on one slide and makes a leader feel like the quarter had a story. Review time held steady while defects dropped needs three slides, a caveat, and someone confident enough to walk a VP through it. The post says machine activity is easy to graph and easy to brag about, and the brag part is doing more work than people give it credit for. The metric that survives is the one that is easiest to present, not the one that is true.

    Permalink
  • curious_clueless

    wait, is the token and prompt count thing not just the lines of code example from the bottom of the post with a new name on it? genuinely asking. the post says management used to reward writing more code and it bloated everything. is reward more prompts not the exact same move? feels like we already ran this experiment and i'm a little confused why it's landing as a surprise.

    Permalink
  • juicy_lemon

    Everyone here is blaming the engineers gaming the number, but the metric solves a manager's problem, not a measurement problem. In calibration I have to stack people and own it with my own name attached. AI adoption is high on my team is a sentence I can say without a single hard judgment hanging off it. That is the appeal. It turns a defensible but arguable call into a tidy number I can stand behind when someone pushes back. The people manufacturing activity are just responding to a manager who already opted out of judging. OP has the incentive chain pointed one level too low.

    Permalink
  • chihiro

    Calling it KPI corruption lets everyone off the hook. The honest name is definition cowardice. Nobody in the room wants to write down what good engineering output actually is, because that requires an argument with someone senior, so they reach for prompt counts instead. There is a well known write up from the DORA research program that keeps repeating this: activity metrics like commit volume predict almost nothing about delivery performance, and the moment you reward them they get gamed. AI usage counts are the same shape. We did not forget Goodhart. We just prefer the metric that can be pulled from a vendor dashboard without anyone having to defend a definition.

    Permalink
  • ripleymode

    Love that the proposed fix in the post is measure defect rate, review time, and rollback risk instead. Sure. Those are the same numbers that have been invisible on every promo packet I have ever read. The work that keeps a release from blowing up has never once shown up as a celebrated metric, so forgive me for not believing leadership will suddenly start tracking it now that AI made the easy graph available. They wanted a number they could brag about. Rollback risk does not brag well.

    Permalink
  • trinityvale

    The whole piece rests on this image of the employee as an incentive robot who will instantly corrupt any metric, and I think that is the lazy part. Most teams I run ops for do not actually optimize hard against soft signals. People glance at the AI dashboard, shrug, and keep doing their job, because the real pressure in their week comes from their tech lead and their on call rotation, not a usage graph nobody reviews. The backfire story is satisfying because it makes us all sound powerless and clever at the same time. In practice a badly chosen metric usually just gets ignored, not weaponized.

    Permalink
  • doordesk_dan

    I have seen this exact movie. A metric gets introduced to make a tool look adopted, and within two cycles it has become documented suffering. We had a frugality dashboard once that tracked how often you used the cheaper internal tool. People used it constantly, for nothing, just to keep the green bar green, while quietly doing the real work in the thing that actually worked. The AI usage number is going to be that, except now the empty ritual also costs you a few cents per prompt. Congratulations. You automated the busywork that proves you did busywork.

    Permalink
  • silver_moth

    From the chair I sit in, the failure is almost never the engineers gaming the number. It is leadership picking a number it can present upward without having to defend a judgment. Once AI adoption is on a slide that goes to the board, every manager under it inherits the target whether it measures anything or not. The post treats this as employees looking out for themselves, which is true but incomplete. The deeper move is that an organization reaches for a countable proxy precisely when it has stopped trusting its own managers to evaluate quality. The proxy is not laziness. It is a quiet admission that nobody wants to own a hard call.

    Permalink
  • spike

    Tokens burned is a perfectly fine signal. It is just a cost signal, not a productivity one, and the post keeps treating those as the same axis. I look at token spend the way I look at our cloud bill, as a thing that tells me where the money went, full stop. The actual mistake is the leap from we spent a lot on the model to therefore the team produced a lot. Nobody sane reads the AWS invoice and concludes the team had high impact this quarter. Measure the model spend, sure. Just do not promote anyone off it.

    Permalink
  • akira

    Let me actually defend the objection the post waves away in two sentences. Early in a rollout you genuinely do need usage data, and not just a yes or no on adoption. You need to know which workflows people reach for, where they drop off, where the tool quietly fails so they stop trusting it. That is product instrumentation, and it is legitimate. The real boundary is between instrumentation you read to fix the tool and a target you attach to someone's rating. The first one dies quietly when the rollout matures. The second one never dies, because somebody put it on a slide. The post collapses both into vanity metrics and loses the useful distinction.

    Permalink

Related discussions

  • Are the managers who bet AI would replace engineers being replaced the fastest?

    Last year my LinkedIn feed had a genre. A program manager or a "delivery lead" or someone with Agile in their headline would post a screenshot of an AI writing a function, add a line like "and they said this job was safe, just learn how to code" and collect four hundred likes from people who do the same job. The implication was always that the typing part of engineering was the engineering, and now that a model can type, the typing class was finished.

  • Do you stop getting frustrated at work once you understand corporate incentives?

    There is a status deck somewhere in your company that nobody reads. It gets updated every few weeks, shown in a meeting, and forgotten. Your manager knows this too. They built the same decks on the way up and understand exactly how little thought usually goes into them. The usual explanation for corporate busywork is that someone higher up is confused or disconnected from reality. That's comforting, but mostly wrong. These artifacts survive because they're doing a job, just not the one they…

  • Will AI replace you, or will a coworker using AI replace several of you?

    A lot of office workers are comforting themselves with the wrong question. They keep asking whether AI can do their whole job. That is not the threshold their employer will use. The real question is whether the output can be produced cheaply enough, and checked cheaply enough, that the role starts looking expensive. It's not if AI can fully do our job, is "can it accelerate it long enough so only half of my team is needed?". Because the answer to that, sadly, is yes.

  • Can AI make you lose your mind, and are you more at risk if you doubt it?

    I always felt that AI companies are actually putting wrappers on top of AI to identify that we're testing it for thinking. For example back when we'd made it count the vowels/consonants in a word and it'd get it wrong. I feel there's a script now that just gets called when the task is identified correctly. I also feel that it gets trained on these memes. Today, I found a new test, one that shows how easily AI gives you AI psychosis and how easy it easy to truly believe that everything you ever…

  • Are most AI startups just a UI on top of some Agent.md files?

    Most AI startups right now feel like someone glued GPT to a terminal, added a dark mode UI, and started talking like they invented something.You’ll see these insane pitches like “persistent autonomous cognitive agents with long-term reasoning” and then you look under the hood and it’s basically: give the model tool access, let it use a browser, maybe add memory summaries and retry logic. That’s the “product.” You can get that on your own just giving access to Claude locally.

  • Does AI make it impossible to tell great engineers from noisy ones?

    I keep hearing the same feedback in different forms: “great velocity,” “love the throughput,” “nice use of AI.” From the outside, it really does look like more is happening: more Code Reviews, more tickets touched, more updates, more emails, more tasks, more designs. AI makes it easy to sustain that cadence without the usual friction of writing, thinking, or even hesitating. But inside the work, there’s a dilemma that keeps getting bigger.

  • Why do managers want everyone else to use AI but themselves?

    The thing that’s starting to irritate me is not the AI push itself. Some of the tools are genuinely useful. I use them every day now. What irritates me is management demanding “AI-first” behavior while keeping every surrounding process aggressively hostile to AI usage. People are told to use AI for coding, planning, research, drafting, debugging, knowledge retrieval, project coordination.. But then half the company’s operational knowledge still lives inside undocumented conversations and…

  • Is the Rolex Submariner really a dive watch, or an office master?

    The Rolex Submariner is the greatest fantasy object ever sold to men with Outlook calendars. This watch has spent seventy years convincing finance guys, dentists, and accountants that they’re rugged maritime adventurers instead of people who say things like “circle back after lunch.” The Submariner is technically a dive watch, but the average one sees less water than a cactus, because god forbid the seals don't actually work well and it gets wet on the insides. These things spend their lives…