Home / Blog / Cost Per Delivered Unit

Day 18

Cost Per Delivered Unit: Using AI to Change Your Margin Structure, Not Your Marketing

Everyone has added AI to the website. Almost nobody has moved the number that decides whether the company is worth running.

By Adrian Dunkley11 min readOperations

There is a version of adopting AI that changes nothing: tools get bought, everyone reports feeling faster, and twelve months later the margin is identical. There is another version where the cost of producing the thing you sell falls by 40 percent and stays there, and every strategic option in the business opens at once.

The difference is one metric. Cost per delivered unit: what it costs you to produce one of whatever a customer pays for. One audit. One campaign. One resolved ticket. One processed claim. Most service and operations businesses have never calculated it, which is why they cannot tell whether any of this is working.

Feeling faster is not a business result. Producing the same unit for less money is.

Start by naming the unit precisely enough that you could count it on a whiteboard. "Consulting" is not a unit. "One completed vendor risk assessment" is. "Support" is not a unit. "One resolved ticket at first contact" is. If you cannot name the unit, you cannot price it, staff it, or improve it, and every efficiency conversation in your company will be about vibes.

56%faster task completion, Copilot randomised trial
40%faster professional writing tasks, Noy and Zhang, Science
14% / 34%support productivity gain: average, and for novices

Those are the credible numbers, from controlled studies rather than vendor decks. Peng and colleagues ran a randomised trial on GitHub Copilot and found a defined programming task completed roughly 56 percent faster. Noy and Zhang, publishing in Science, found mid-level professional writing tasks completed about 40 percent faster with quality rated higher. Brynjolfsson, Li and Raymond studied thousands of customer support agents and found roughly 14 percent more issues resolved per hour on average, rising to about 34 percent for the least experienced staff.

Read the shape rather than the headline. Gains are large on bounded, well-specified tasks, and they concentrate among less experienced workers. That tells you exactly where to apply this in your business, and where not to bother.

Measure the unit before you change anything

Take twenty recently delivered units. For each one, list every step, who touched it, and how long they spent. Multiply time by loaded hourly cost, which is salary plus employment costs plus overhead, divided by realistic productive hours, not contracted hours. Add software and infrastructure consumed. Add rework: the time spent fixing what came back.

You will discover two things immediately. The distribution is wider than you thought, with the worst unit costing three or four times the best one. And a step nobody talks about, usually a handover or a wait, consumes more than the work itself.

The loop · from guesswork to a number that moves
1

Name the unit

One countable thing a customer pays for. Specific enough to tally on a wall.

2

Time twenty of them

Step by step, with loaded cost and rework included. Find the widest step.

3

Automate one step

High volume, low variance, checkable in seconds, cheap when wrong.

4

Keep the reviewer

Measure error rate against the human baseline on the same sample.

5

Bank the saving

More units, fewer hours, or a higher price. Decide which before you start.

Five steps, run per unit type, repeated quarterly. Step five is the one companies skip, which is why so many report enthusiasm and no margin change.

Run your own numbers

Use minutes per unit and your real loaded hourly cost. Review time is not optional, so include it honestly: automation moves work from doing to checking, it does not delete the human.

Interactive · cost per delivered unit, before and after
$26.25 new cost per unit (was $67.50)
$16,500 monthly cost removed at current volume
2.6x capacity per person at the same headcount

Under 10 percent saved. Not worth the change management. Pick a different step or a different unit. 10 to 30 percent. Real but modest. Bank it as capacity, not as headcount, and move to the next step. 30 to 60 percent. This is the band the controlled studies report on bounded tasks. Decide now how you bank it. Over 60 percent. Verify the review time is honest. If it holds, your pricing model and your capacity plan both need rewriting this quarter.

Cost per unit is minutes divided by 60, times loaded hourly cost. At 90 minutes and 45 an hour, a unit costs 67.50. Cut to 35 minutes including review and it costs 26.25, removing 16,500 a month at 400 units and giving each person 2.6 times the capacity. That last number is what you sell, hire or price against.

The capacity multiple is the number to take into your planning. At 2.6 times, the delivery team you have can serve a pipeline you previously needed to double headcount for. That is either a hiring freeze, a growth plan, or a price cut that takes market share, and you should say out loud which one you are choosing.

The savings you never bank

Time saved is not money saved. It becomes money in exactly three ways, and if you do not pick one, the minutes disappear into the general expansion of work.

Sell more units with the same team. The best outcome when demand exists. Requires your pipeline to be the constraint, not your delivery.

Reduce the hours you pay for. Fewer contractors, less overtime, a role you do not backfill. Immediate and visible in the P&L, and the one that needs the most honesty with your team about what is happening and why.

Hold price while cost falls. Margin expansion, invisible to customers, available immediately if you priced on value rather than hours. If you bill hourly, note the trap: getting twice as fast cuts your revenue in half. That is the strongest argument for fixed-price or per-unit billing in a professional services business, and it is now an urgent one rather than a philosophical one.

Pick the one you want before the project starts, and write down the number that should move. "Contractor spend down 40 percent by October" is a decision. "Improved efficiency" is a mood.

Do this now: the twenty unit teardown

Block three hours this week. Take your last twenty delivered units and build a table: unit, total minutes, minutes per step, who touched it, rework minutes, direct software cost. Sort by total cost. Look at the top three most expensive and the three cheapest, and write one sentence each on what made the difference. In almost every business I have done this with, the answer is not skill. It is whether the input arrived complete. Which means the highest-return automation is usually not in the work itself, it is at the intake: validating, chasing and structuring what the customer gives you before anyone starts. Fix intake and the whole distribution tightens.

Where automation earns its place, and where it burns you

Sort every step in your unit against two questions: how expensive is an error, and how quickly can you detect one?

Cheap error, fast detection: automate immediately. Drafting, classification, extraction, formatting, first-pass research, summarising a call. A person reads the output in seconds and knows if it is wrong.

Expensive error, fast detection: automate with a mandatory reviewer and a measured error rate. Most professional work sits here, and the reviewer is the product, not a compromise.

Cheap error, slow detection: instrument before you automate, because you will not notice the drift for months. Anything that quietly feeds a report someone reads quarterly.

Expensive error, slow detection: leave it alone for now. Final financial sign-off, legal filings, clinical decisions. The saving is small relative to the tail risk, and the tail risk is the kind that ends companies rather than quarters.

What controlled studies actually measured
Coding task, Copilot RCT56%
Professional writing, Science40%
Support agents, novices34%
Support agents, average14%

Measured productivity gains from published studies, not vendor claims. Note the spread: a bounded task with immediate verification gained 56 percent, while an average across a whole support function gained 14. The closer a task is to "one clear output, checkable now," the larger the gain.

The gap between 56 and 14 is the most useful thing in that chart. It is the difference between measuring a task and measuring a job. Your business is made of jobs, so plan for the lower number and be pleased when a specific step lands nearer the higher one.

The quality question, answered with a sample

Someone on your team will say the output is not good enough. They may be right, and the argument cannot be settled by opinion. Settle it with a sample.

Take fifty units. Produce each one both ways. Have a reviewer who does not know which is which score them against your actual acceptance criteria: factual accuracy, completeness, tone, adherence to the template. Count errors, and count how long each took. Now you have an error rate and a time, for both methods, on the same work.

Two outcomes are useful. If the automated path has a higher error rate but takes a third of the time, the question becomes whether review closes the gap for less than the time saved, which is arithmetic rather than argument. If the error rate is the same, you have just retired a debate that would otherwise have run for six months.

Keep the sample. Re-run it every time you change models or prompts, because a silent regression in output quality is the single most expensive failure mode in an automated pipeline, and the only defence is a fixed test set you trust.

The takeaway

  • Name your delivered unit precisely, then measure what twenty of them actually cost, rework included.
  • Expect large gains on bounded, checkable tasks and much smaller ones across whole jobs. Plan for the smaller number.
  • Automate where errors are cheap or fast to detect. Leave expensive, slow-to-detect steps alone.
  • Decide in advance how you bank the saving: more units, fewer hours, or held price. Otherwise it evaporates.
  • If you bill by the hour, speed cuts your revenue. Move to fixed or per-unit pricing before you automate.
  • Settle quality arguments with a fifty-unit blind sample, and keep it as a regression test.

Frequently asked questions

What is cost per delivered unit?

The full cost of producing one of whatever you sell: one audit, one campaign, one resolved ticket. It includes loaded people cost, software and infrastructure consumed, and rework. It decides whether growth improves or destroys your margin, and most service businesses have never calculated it.

How much does AI actually improve productivity?

Controlled studies: a Copilot randomised trial found a coding task about 56 percent faster; Noy and Zhang in Science found professional writing roughly 40 percent faster at higher rated quality; Brynjolfsson, Li and Raymond found support agents resolved about 14 percent more issues per hour, near 34 percent for novices. Large gains on bounded tasks, concentrated among less experienced staff.

Why do AI time savings not show up in profit?

Saved minutes only become money if you sell more units with the same team, pay for fewer hours, or hold price while cost falls. If you pick none of those, the time is absorbed by other work. Choose one in advance and measure that number, not the minutes.

Which tasks should you automate first?

High volume, low variance, checkable in seconds, cheap when wrong: drafting, classification, extraction, first-pass research, formatting. Avoid steps where errors are expensive and slow to detect, such as final legal or financial sign-off, until you have measured your error rate against a human baseline on the same sample.

The tools were adopted. The margin never moved.

Kill My Startup is about the difference between activity that looks like progress and the numbers that decide whether you survive.

Buy on Amazon →

Sources

  1. Peng, Kalliamvakou, Cihon and Demirer, randomised controlled trial of GitHub Copilot on developer task completion time.
  2. Noy, S. and Zhang, W., "Experimental evidence on the productivity effects of generative artificial intelligence," Science.
  3. Brynjolfsson, E., Li, D. and Raymond, L., "Generative AI at Work," on customer support agent productivity and the distribution of gains.
  4. Standard cost accounting definitions of loaded labour cost, rework and cost per unit.