Measuring Time Saved Without Lying to Yourself
Ask any content marketer who has been using AI for more than a fortnight how much time it saves them, and you will get a number. Usually it is round, usually it is large, and usually it is wrong. “Cuts my first drafts from four hours to forty minutes.” “Probably 60% faster.” “I reckon it gives me a day a week back.”
These numbers aren’t lies. They’re honest reports of a badly-scoped measurement. The person is timing the part of the work that changed and ignoring the part that got longer.
Where the missing time goes
When you draft with Claude or ChatGPT, the visible work collapses. A 1,400-word first draft that used to take you three hours arrives in six minutes. That gap is real, it’s dramatic, and it’s the number your brain files away as “the time saved.”
What your brain doesn’t file is everything downstream. The fabricated stat you spent eleven minutes chasing before deciding it didn’t exist. The three sections that read fine in isolation but say the same thing in different words, so you cut two and rewrote the third from scratch. The client’s SME flagging that your explanation of their pricing model was subtly, confidently wrong, and the forty-minute call to fix it. The bit where you rewrote the intro four times because the AI version was competent and dead.
None of that feels like production time. It feels like editing, which you were always going to do, so it gets charged to a different mental account. That’s the whole mechanism: AI shifts effort from a phase you time to phases you don’t. Drafting is legible and bounded. Verification is diffuse, interruptible, and spread over days, so it never gets counted.
There’s a second mechanism stacked on top. Editing AI prose is not the same activity as editing your own. When you edit your own draft you’re refining something you already believe. When you edit AI output you’re doing forensic work: checking whether each claim is true, whether the structure actually argues anything, whether the voice is yours. That’s a different cognitive load at roughly the same word count, and it produces more passes, not fewer.
Add the two together and a “45 minutes instead of 3 hours” claim routinely nets out at 2h10 against a 3h baseline. Still a win. Not the win anyone reported.
The lightweight method
You don’t need a time-tracking discipline. You need one that survives a busy week, which means it has to be dumber than you think is acceptable.
Here’s the whole thing. Four phase buckets, one row per asset, timer running in Toggl Track with four hotkeyed entries (Clockify and Harvest work identically; a spreadsheet with start/stop columns works too, it’s just leakier).
The buckets:
| Bucket | What goes in it |
|---|---|
BRIEF | Research, outline, prompt construction, sourcing. Everything before you have a draft. |
DRAFT | Generating and assembling the first complete version, whether you typed it or prompted it. |
EDIT | Structural rewrites, voice passes, cuts, line edits. |
VERIFY | Fact-checking, link-checking, chasing claims, SME review time (theirs and yours), and post-publish corrections. |
The critical rule: VERIFY stays open until the asset has been live for seven days. Every correction, every “actually that figure is from 2019,” every Slack thread from the product lead goes on that asset’s clock. This is the rework tail, and it’s the entire reason self-reported figures are inflated. Close your books too early and you measure the good half.
Second rule: log SME and reviewer minutes, not just yours. If your AI draft costs your head of product twenty minutes of correction that your hand-written draft wouldn’t have, the programme spent twenty minutes. Whose timesheet it lands on is an accounting question, not a productivity one.
Getting a baseline you can defend
You cannot measure savings without a comparison, and “how long it used to take” is not a comparison. It’s a memory, and memories of pre-AI workflows are compressed and flattering.
Run a paired batch. Take ten assets of the same type and rough complexity, from the same brief template: five human-first, five AI-first, alternating so your energy levels distribute across both arms. Same writer where possible. Two weeks of work for most small teams, and it’s the only honest baseline you’ll get.
If you genuinely can’t spare that, the fallback is a rolling baseline: log every asset in all four buckets for six weeks regardless of method, then segment retrospectively. Noisier, no setup cost, and it beats guessing.
A worked example
Here’s a real shape of result from a three-person in-house team producing mid-funnel B2B SaaS explainers, roughly 1,400 words, one SME review each. Ten assets, paired.
HUMAN-FIRST (n=5, mean minutes)
BRIEF 52
DRAFT 168
EDIT 61
VERIFY 34 (incl. 12 min SME, 4 min post-publish)
TOTAL 315 = 5h 15m
AI-FIRST (n=5, mean minutes)
BRIEF 71 (+19 prompt construction, source gathering)
DRAFT 41 (-127 the number everyone quotes)
EDIT 104 (+43 structural rewrites, voice pass)
VERIFY 79 (+45 incl. 31 min SME, 17 min post-publish)
TOTAL 295 = 4h 55m
NET SAVING: 20 minutes per asset (6.3%)
SELF-REPORTED SAVING (survey, same team, same fortnight): "about 60%"
Sixty per cent versus six. That gap is the point of this article.
Look at where it comes from. The DRAFT line is exactly as good as everyone says: 168 minutes down to 41, a 75% cut. If you only time drafting, you report 75% and you’re not lying, you’re just answering a smaller question than the one you were asked. Meanwhile BRIEF went up (good prompting is real work), EDIT went up 70%, and VERIFY more than doubled, with the SME burden rising from 12 minutes to 31 because reviewers now find substantive errors instead of typos.
The 17 minutes of post-publish correction in the AI arm against 4 in the human arm is the tail you’d miss by closing the clock at publication. Four of the five AI-first assets needed a post-live fix. One had a misattributed quote that sat on the site for nine days.
Now the awkward part: 6.3% is not nothing, but it’s nowhere near enough to justify the tooling spend, the workflow rebuild and the internal evangelism. The team’s honest conclusion was that AI-first drafting was roughly break-even for this asset type, and they moved it to where the numbers were better.
Where the numbers actually were
Same team, same method, applied to two other jobs:
Repurposing a published explainer into five LinkedIn posts and a newsletter section. Human-first: 94 minutes. AI-first: 38 minutes. VERIFY barely moved (12 to 16 minutes) because every claim was already checked in the source asset. Net saving 56 minutes, 60%. The rework tail stays flat when the AI isn’t allowed to originate facts.
Drafting the brief itself: SERP analysis, angle options, outline, internal link suggestions. 52 minutes to 23. VERIFY didn’t apply at all because a bad outline gets fixed in the same session, not nine days later on a live page.
That’s the pattern the four-bucket method surfaces and gut feel doesn’t: AI saves the most time on tasks where verification cost is near zero, and the least on tasks where it originates claims. Transformation is cheap. Origination is expensive. You cannot see this distinction if you only time the draft, which is why so many teams push AI hardest at exactly the task where it pays worst.
What to report upward
Give your director or your client one table per asset type, with the four buckets, the paired means, and the net figure. Add the self-reported number next to it and label it honestly as perceived. That contrast does more for your credibility than any headline percentage, because everyone in the room has a private suspicion that the 60% figures are soft, and you’ll be the one who checked.
Two practical notes on the reporting. State your n and don’t round up: “20 minutes per asset across five paired assets” is a defensible sentence, “about 6%” invites a demand for precision you don’t have. And express savings in minutes per asset before you express them in days per month, because the multiplication is where credibility goes to die. Five assets a week at 20 minutes is 1.7 hours a month, not “a day a week.”
Net time is also only half the ledger. Twenty minutes saved on an asset that underperforms is worth less than forty minutes spent on one that ranks, and the four-bucket method says nothing about output quality. Pair it with the performance side: our guide to measuring AI content performance and ROI covers how to connect per-asset cost to per-asset outcome so you’re not optimising for speed on content nobody reads.
The bits that will go wrong
You’ll forget to start the timer. Everyone does. Accept 80% capture and note which assets are partial rather than binning the whole exercise. A dataset with three flagged rows beats an abandoned spreadsheet.
Your AI-first numbers will improve over six weeks, and you won’t know why. Could be genuine skill gain at prompting, could be that you’ve quietly stopped verifying as carefully. Track VERIFY as a percentage of total per asset over time. If it drops while your published-correction count holds steady, that’s real. If corrections are rising, you’ve traded measurement for complacency.
Someone will argue the baseline was inflated because the human-first drafts were written by your most senior person on a bad week. Fair challenge, and the answer is to name the writer on every row from day one so the objection is answerable with data rather than defensiveness.
The rework tail is the first thing you’ll want to drop, because waiting seven days to close a clock is annoying and the corrections arrive in Slack rather than your tracker. Put a recurring Friday reminder on it: fifteen minutes, sweep the week’s live assets, log any fixes. That fifteen minutes is what separates a number you can defend from a number you made up.
Run it for six weeks on one asset type. You’ll either find the savings are smaller than you claimed, in which case you’ve stopped a bad bet early, or you’ll find they’re real and you can prove it to someone holding a budget. Both outcomes are worth more than the 60% you currently believe.