Home → Leadership
Can You Measure Performance With OKRs?
End of year. A dashboard on screen: 72% complete. Everyone is looking at it. And nobody in the room knows whether that number is good or bad.
- OKR is an alignment tool, not a measuring instrument. Confusing the two breaks both.
- Tie a bonus to a target and people shrink the target. That is not cunning; it is the behaviour the system taught them.
- Story points are not a performance unit. The day you start measuring them, they start inflating.
- Performance is read from three sources: outcome, contribution, growth. And it is written down through the year, not remembered at the end.
What were OKRs actually invented for?
One job: getting everyone to look in the same direction. At the start of a quarter the company says "these three things matter now"; teams connect their work to that sentence. The point is to stop ten teams rowing in ten directions.
To do that it encourages ambitious targets. The classic advice is that hitting 70% is good, because if you always hit 100% you set the bar too low.
Now notice: that system rests on failure being safe. And we destroy that safety with our own hands.
What happens the moment you attach a bonus
Say you tie OKR achievement to bonuses. Targets get written in January. The conversation that actually happens is this:
— "Shall we take response time from 400 to 150 ms?"
— "Risky. We would need a cache layer, might not make it. Let us write 300,
we will hit that. If we do better it looks great anyway."
Nobody in that conversation is acting in bad faith. Both are answering the system correctly. You said "there is money if you hit it", so they picked a target they will hit. Result: the dashboard is green and the company attempted less.
Three examples from the field
1. "Tickets closed"
A support team's performance indicator became "tickets closed". Three months later the number was up 40%; everyone was pleased. Then someone read the tickets.
What used to be resolved in one ticket — "the user cannot log in, and also cannot see the report" — was now opened as two separate tickets and closed twice. Nobody lied, nobody slacked. They simply learned how the counter goes up.
2. Story point inflation
One team was given a "raise velocity" target. Six months later velocity really had gone up: from 34 points per sprint to 52. The number of features delivered stayed the same.
The reason is simple: work that used to be called 3 points was now called 5. Nobody conspired; it just got easier to say "this is a bit harder actually" in estimation and harder to object. A story point is a unit of estimation; the moment you load performance onto it, it stops being one.
3. The team that hit 100%
At the end of one quarter a team hit every target at 100%. The first instinct was to congratulate them. Then I looked at the targets: three were work already underway, and one was nearly finished when it was written.
So the team did not work badly; they wrote the targets backwards. And the reason they did is that the previous quarter, a team that set an ambitious target and missed it got told off in a meeting. People do whatever the system teaches.
So how do you assess performance?
I am not saying do not measure — if you do not, reviews reward whoever is loudest, which is the worst outcome. But reducing it to one number does not work either. What works in practice is reading three sources together:
| Source | The question | Where evidence comes from |
|---|---|---|
| Outcome | What shipped, and who did it help? | Things live in production, problems solved, measurable impact |
| Contribution | Did they improve the people around them? | Code reviews, mentoring, documentation, peers' views |
| Growth | Where are they compared to a year ago? | Difficulty of work taken on, increased independence, decisions made |
Notice the middle column: none of them asks "how many points did they do". Because the most valuable engineers are often the ones with low personal numbers — because they unblocked someone else, fixed the pipeline, kept the infrastructure standing.
And the most important rule: write it down through the year
The weakest link in reviews is memory. Sit down in December and try to recall the year and you will recall the last two months. That is not dishonesty, it is how brains work.
The fix is boring but effective: ten minutes a month, two or three lines for each person on the team. What they did, what difference it made, where they struggled. By December you have twelve months of real record and the conversation does not start with "my impression is".
If people at the same level are assessed by different managers, the ratings measure the managers' generosity, not the people's performance. The generous manager's team gets promoted, the strict one's team gets bitter.
The cure is a calibration session: managers in one room, people at the same level placed side by side, each rating defended out loud. It is an uncomfortable meeting. That is the price of fairness.
Should we drop OKRs entirely?
No. Just let them do their own job and not someone else's:
- Showing in one place what matters this quarter
- Stopping teams from working in ignorance of each other
- Answering "why are we doing this?"
- Making ambitious targets safe
- Bonus and promotion decisions
- Individual ratings
- Pitting teams against each other
- Producing reports for leadership
A practical dividing rule: OKRs belong to the team, performance belongs to the person. Keep OKRs at team level and do not write individual ones. What you discuss individually is not a percentage but how that person contributed to it.
Back to that 72%
The percentage on its own says nothing, because it squeezes two different stories into the same number:
- A very ambitious target was set, the team stretched, 72% is excellent.
- A comfortable target was set and priorities shifted mid-quarter, 72% is poor.
So at every OKR close, the question before the percentage is: "What did we learn this quarter?" An OKR cycle with no answer to that has only filled in a spreadsheet.
Checklist
- Does the OKR percentage feed directly into bonus or promotion? (If so, separate them.)
- Do I have notes from through the year, or am I recalling the last two months?
- Does the assessment cover outcome, contribution and growth?
- Was there a calibration session for people at the same level?
- Will any of my criticisms be a surprise to the person?
- Have I thought about how each number I track could be gamed?
Conclusion
OKRs do not measure performance; they state what the company cares about. Turning that into a ruler costs you two things at once: targets shrink and reviews become untrustworthy.
Performance does not fit into a single number anyway. Every time you try to squeeze it in, the team learns how to grow that number — usually faster than you do.