Measure Productivity by Output, Not Activity: A GTD Guide

The only productivity number worth keeping is the count of loops you closed this week. Everything else measures intent.
This piece shows you how to measure productivity by output, not activity, inside a working GTD system rather than in the abstract. The conclusion, up front: define each project’s done-condition as a past-tense sentence someone else could verify, count four things during your weekly review (projects closed, deliverables shipped, decisions settled, loops reopened), and turn off every completed-task total and streak your apps are showing you. That is the whole method. The rest is the detail that stops it going wrong.
Ticking Off Tasks Is Not Producing Anything
Next actions in GTD are inputs by design. A next action is the smallest physical, visible step you can take, which makes it a brilliant unit for getting started and a useless unit for measuring delivery. When you count completed next actions you are counting how many times you intended to move something forward. Intent is not output.
The measurable unit is the closed loop. David Allen’s framing of open loops is the sharpest test available: a loop is anything you have committed to that is not yet where it belongs. It closes when the result leaves your hands. Sent, filed, published, signed, handed over, decided and communicated. Not “nearly ready”. Not “drafted and sitting in a folder”.
Take a freelancer whose task app proudly shows 312 completed tasks across a month. Look at what actually left their hands: one proposal sent, two articles filed with the client, one invoice raised. Four outputs. The 312 is motion, and the gap between the two numbers is where the feeling of being busy and broke lives at the same time.
False signals get treated as productivity every day:
- Inbox zero. A processed inbox means you have made decisions, which is good. It does not mean anything shipped.
- Completed-task counts. Rewards granularity, punishes people who write bigger actions.
- App streaks and productivity scores. Measure app usage. Nothing else.
- Hours logged. Measures availability.
- Meetings attended. Measures presence, and often measures other people’s loops.
Then there is the rework trap. Fast activity frequently opens more loops than it closes: the hurried reply that triggers three clarifying emails, the deliverable shipped at 80 per cent that comes back with comments, the decision made without the one person who then reverses it. High activity with negative net loop closure is entirely possible, and it feels like the busiest week of your year.
Define Your Outputs Before You Measure Anything
You cannot count what you have not defined. GTD’s project start-up gives you a free test: write the successful outcome as a sentence in the past tense that a colleague could verify without asking you how it went.
“New pricing page live.””Q3 budget approved by the board.””Supplier contract signed and filed.”Those are verifiable. If you cannot write that sentence, the project is not measurable, and in my experience that is usually why it has been sitting on your list for five months. It stalls because it has no finish line, not because you lack discipline. Motivation gets blamed for a definition problem.
Three kinds of output, and why the distinction matters
- Deliverables: things handed over. A report issued, code released, an invoice sent, a client call delivered.
- Decisions: things settled and communicated, so nobody has to think about them again. Choosing the supplier counts only when the other two have been told.
- Artefacts: things that now exist and did not before. A documented process, a template, a tested backup, a hiring brief.
Vague projects produce none of the above. “Sort out the website”is a permanently open loop pretending to be a project. Split it and you get three countable outputs: new pricing page live, contact form tested and receiving enquiries, old blog posts redirected. Same work. Three finish lines instead of nought.
Try this before you measure anything. Open your project list and read each entry. Mark every one where you cannot name the thing that will exist, move or be settled when it is done. Those marked projects are your real backlog, and they are the reason your output count will look thin in week one.
The Weekly Output Count, Bolted Onto Your Review
Do not build a reporting habit. Build a column. The weekly review already exists in your system, you are already looking at your project list, and that is the natural moment to count what closed. A separate tracking ritual fails for the same reason a note app with a daily admin tax in tags and statuses fails: anything that costs maintenance time gets quietly abandoned inside a fortnight.
Record four numbers, nothing more:
- Projects closed. Done-condition met and verified.
- Deliverables shipped. Things that physically left your hands this week.
- Decisions settled. Decided and communicated.
- Loops reopened. Work that came back, was rejected, or needed redoing.
Keep it to one page. Analogue wins here and it is not nostalgia: a single index card or one notebook page holds a week’s tally with no syncing, no configuration and no temptation to spend Sunday evening building a dashboard instead of doing the work. If you want something pre-ruled, the printable hPDA cards on this site take a four-column tally without modification. A digital note is fine too, as long as it stays one note.
Now the discipline that most people skip. Run it for four weeks before you conclude anything. Four rows, four columns, no interventions, no new system, no changing your project list to make the numbers look better. One good week and one wiped-out week tell you nothing. A four-week baseline tells you your realistic rate, and a realistic rate is what makes next quarter’s commitments honest.
Activity Metrics Worth Keeping, and the Ones to Switch Off
Not all activity data is rubbish. A few leading indicators protect output rather than impersonate it.
Time to captured. How many seconds from having the thought to it being safely in your system, with your phone locked and your hands full. Under five seconds is a practical working benchmark rather than an official rule. Slow capture means lost commitments, and lost commitments are lost output later. If you want to see how much that varies by tool, the comparison of note apps scored on capture speed covers it properly.
Cycle time. Date captured to date the loop closed. It shows you where work stalls without counting effort at all, and it is brutally informative on projects you thought were “nearly done”six weeks ago.
Waiting-for list length. The honest measure of work you cannot finish alone. A growing waiting-for list with flat output is not a personal performance problem, and it is the single most useful thing to take into a conversation with a manager or a client.
What to delete: completed-task totals, streaks, badges, productivity scores, screen-time dashboards. Task apps push these hard because churn looks like engagement. Turning those displays off is a measurement decision, not a cosmetic one, and if your app lets you build custom views, spend that effort on views that show stalled projects instead. The custom OmniFocus perspectives guide is one way to do that.
Keystroke logging and presence tracking deserve a blunter verdict. Skip them. Beyond the fact that they measure nothing you can ship, UK employers monitoring staff have obligations under data protection law, and the ICO’s guidance on monitoring workers sets out that monitoring must be lawful, fair and proportionate. Check the current version before anyone installs anything.
Measuring Output When the Work Is Invisible or Shared
The objection to output counting is always the same: my work does not produce widgets. Fair. It still produces something.
Thinking and planning work. Run the natural planning model (purpose, outcome, brainstorm, organise, next actions) and treat the finished plan as a genuine output. A written plan that someone else could act on is an artefact. A decision made and communicated is a deliverable. The David Allen Company’s own material on the methodology is worth reading if you have only ever met GTD through blog summaries.
Teams. Count at the boundary where work leaves the team, never per person. Releases shipped, client reports issued, cases resolved. Per-person ticket counts reward people for picking easy tickets and punish the person who spent three days unblocking everyone else. If you want individual data, use cycle time on the team’s shared work, not volume on individuals.
Email-heavy and support roles. Messages handled is the classic vanity number. Do the GTD inbox pass properly, deciding for each item whether to do it now, delegate it, defer it into your system, file it as reference or delete it, then measure resolved threads and reference filed. Two hundred messages touched and forty threads resolved is a real week. Two hundred messages touched and eight threads resolved is a warning. Your mail client affects how fast that pass runs more than most people expect.
Long projects. A three-month build with one output at the end gives you eleven weeks of zeroes and one week of triumph, which is useless as feedback and corrosive when you are already feeling overwhelmed at work. Define interim outputs at roughly fortnightly intervals: schema agreed and documented, staging environment live, first module reviewed and signed off. Each one has a verifiable past-tense sentence. Each one counts.
Where Output Measurement Goes Wrong
Gaming the count. Split “quarterly report issued”into nine sub-projects and your tally triples overnight. The done-condition test is the defence: if the outcome sentence is not something a colleague could independently verify as a real result, it is a next action wearing a costume.
Punishing quality. Ship fast, ship broken, count the win, absorb the complaint next week as new work. Record reopened loops honestly in that fourth column and the pattern surfaces immediately. Otherwise a high output count hides a high rework rate, which is precisely the failure that activity metrics get accused of. A week of six shipped and four reopened is worse than a week of three shipped and none.
Ignoring maintenance. Backups tested, dependencies patched, the invoice chased before it went bad. No new artefact, no obvious output, but the absence of it produces very expensive loops later. Keep a short maintenance line on the card rather than pretending it did not happen, and never let the tally become an argument for skipping it.
The honest answer. Sometimes four weeks of data show one or two closed projects a month and nothing hidden. That normally means your projects are too big, not that you are lazy. Split them, rewrite the done-conditions, count again. The numbers usually move within a fortnight, and the feeling of being permanently behind moves with them.
Measure productivity by output, not activity, for one quarter and you will lose the comfort of a large green number in an app. What you get instead is a small, true one, plus a rate you can plan against. Start at your next weekly review with four columns and an honest count of what actually left your hands.
Frequently Asked Questions About Measuring Output Instead of Activity
What does it mean to measure productivity by output rather than activity?
It means counting finished results that left your hands, not the steps you took towards them. In GTD terms, you count closed loops: projects whose done-condition has been met and whose result has been delivered, decided or filed. Tasks ticked, hours worked and emails answered are activity, however satisfying the totals look.
How do I count output if my work is thinking, planning or decisions?
Run the natural planning model and treat the finished plan as the output. A documented plan someone else could act on, a decision made and communicated, or a recommendation written up all pass the done-condition test. Write the outcome as a past-tense sentence (“pricing approach agreed and circulated”) and it becomes countable.
Are completed-task counts in a task app a useful productivity metric?
No. They reward whoever writes the smallest next actions and tell you nothing about delivery. Productivity metrics that track output need a verifiable result behind them, so switch off completed-task totals, streaks and productivity scores and count shipped deliverables instead.
How often should I measure output, and how long before the numbers mean anything?
Weekly, during your existing GTD weekly review, on one page. Give it four weeks before you draw a conclusion, because a single strong or ruined week says nothing about your rate. After a month you have a baseline honest enough to plan commitments around.
Can you measure output for a small team without individual tracking?
Yes, and it works better. Count deliverables at the boundary where work leaves the team, such as releases shipped or client reports issued, rather than tickets closed per person. Add cycle time and the length of the shared waiting-for list, and you can see where work stalls without monitoring anybody’s keystrokes.
- 0 Comment... What do you think?
- Subscribe to RSS
Comments are closed.





