Output vs Outcome
Output is what you produced. Outcome is what changed because of it.
Ten features shipped is output. Users staying longer is an outcome. Forty calls made is output; deals closed is an outcome. Eight hours logged is output, and it is barely even that.
The distinction sounds academic until you notice that almost every metric in use measures output and is treated as though it measured outcome.
Why everyone measures the wrong one
Not stupidity. Output has three properties outcomes do not.
It is countable now. Features, calls, tickets, words, hours. Outcomes arrive weeks or quarters later, if they can be isolated at all.
It is attributable. You can say who produced it. Outcomes are usually caused by several people, market conditions and luck, which makes them useless for evaluating an individual and awkward for evaluating a team.
It is controllable. You can decide to make forty calls. You cannot decide to close ten deals, and being measured on something you do not control is genuinely unfair.
So output metrics survive because they are measurable, attributable and fair, and they persist despite being weakly connected to anything that matters. That is a real trade-off rather than an obvious error.
What goes wrong anyway
Goodhart's law. When a measure becomes a target, it stops being a good measure. Measure tickets closed and tickets get split. Measure lines of code and code gets verbose. Measure hours and hours appear. This is not cynicism about people; it is what happens to any proxy under pressure.
The invisible work disappears. Reviewing someone's work, unblocking a colleague, deciding not to build something, preventing a bad decision in a meeting. All high-value, none of it appears in any output count, so the people doing most of it look least productive.
Thinking scores zero. Every activity-based metric treats the most valuable part of knowledge work as idleness.
Deleting counts as negative. Removing a feature that was hurting the product is an excellent outcome and negative output on most measures.
The useful middle: leading indicators
Pure outcome measurement is impractical — too slow, too confounded — so the working answer is not "measure outcomes instead". It is to pick output measures with a demonstrated link to outcomes you care about.
The DORA metrics are the best-known example in software: deployment frequency, lead time, change failure rate, time to restore. All measurable weekly, all output-ish, and each has been argued to correlate with organisational performance. They work because the link was established rather than assumed.
The test for any metric: can you state the outcome it is a proxy for, and is there evidence for the link? "Hours worked" fails immediately, because nobody can articulate what outcome it predicts.
For measuring yourself
This is where the distinction is most useful, and least fraught, since nobody is being evaluated.
Do not judge a day by output. A day of hard thinking that produces nothing visible can be the most valuable day of the week, and judging by output guarantees you feel bad about it. That mechanism is most of productivity guilt.
Do not judge a day by hours either. Hours are an input, not even output. Eight hours of fragmented attention is not eight hours of anything.
Judge the week by whether the two or three things that mattered moved. A week is the smallest honest unit, and outcomes do not resolve daily.
Track inputs to diagnose, not to score. Where the time went is diagnostic information — it explains why the week went as it did. It is not a grade.
What a time tracker can and cannot tell you

We make a time tracker, so we should be straight about which side of this it sits on.
Cronus measures input. Where your hours went, what you were in, how fragmented it was. That is upstream even of output, and it cannot tell you whether any of it was worth doing.
Anyone selling a tracker as a productivity measure is selling you a proxy for a proxy.
What it is genuinely good for is explaining a result you already have. You know the week went badly. The record tells you your longest unbroken block was 22 minutes and eleven hours went to meetings. That is a cause, and causes are actionable in a way scores are not.
It records automatically on your Mac, with no timer and nothing to log, so the record covers the chaotic weeks as well as the calm ones — and the chaotic weeks are the ones worth explaining.
$6 a month after a three-day trial, and it never captures your screen: window titles only, redacted on your device.
And the reason self-report will not substitute: in a randomised controlled trial, experienced developers estimated a tool had made them 20% faster on real tasks. They were 19% slower — a forty point error about work they had just completed. The study. Your impression of the week is not a measurement of any kind.
If you are measuring other people
Three things worth holding onto.
Never use activity metrics for individuals. Keystrokes, active hours, tickets closed. They are gameable, they punish the invisible work, and they change behaviour toward looking busy. More on why that backfires.
Measure teams on outcomes, individuals on judgement. Outcomes are too confounded to attribute to one person; that is what managers are for.
State the outcome each metric proxies for. If nobody can, the metric is decoration and it is doing harm.
Read next
- Busywork — output with no outcome attached.
- Productivity Guilt — what judging days by output does to you.
- Time Theft — where activity metrics lead.
