Baseline Your Content Performance Before You Introduce AI
Here’s the uncomfortable thing about AI content ROI measurement: almost nobody can do it, because almost nobody wrote down what content cost them before. If you don’t know your cost per asset, your median cycle time and your per-piece performance distribution as of today, then in six months every sentence you say about AI’s impact will be unfalsifiable. “It’s saved us loads of time” cannot be checked. “Quality held up” cannot be checked. “We’re producing three times the output” can be checked, but only on volume, which is the one number that was never the problem.
Unfalsifiable claims aren’t a communication failure. They’re the default outcome of introducing a new production method into a programme with no instrumentation. And they cut both ways: the same missing baseline that lets an over-enthusiastic colleague claim a 60% efficiency gain also leaves you defenceless when a sceptical finance director decides the AI-assisted posts are the reason organic traffic dipped in April, when actually it was a core update and a site migration.
So before the first AI-assisted brief ships, capture twelve numbers. Not forty. Twelve, because a three-person team can actually collect twelve, and a baseline you abandon in week three is worse than none.
What a baseline means when you have 3.5 people, not a data team
A baseline is a dated snapshot of your last twelve months of published assets, at asset level, with cost and timing attached. Asset level is the part people skip. Programme-level dashboards (total organic sessions, total conversions, blended cost) are useless here, because AI will be applied to individual pieces, unevenly, and you need to be able to compare like cohorts later.
The working example through this piece is a UK B2B SaaS team: 3.5 FTE content, one regular freelance writer, 68 assets published in the twelve months to 30 September 2026. Loaded internal rate of £35/hour (£42,000 mid-weight salary, plus 28% for employer NI, pension, software licences and desk, divided by 1,540 genuinely productive hours). Substitute your own rate, but do calculate it. A rate pulled from thin air propagates into every cost figure below.
The twelve numbers
| # | Number | Example baseline | Where it comes from |
|---|---|---|---|
| 1 | Fully loaded cost per published asset | £654 standard post, £1,620 pillar | Hours × loaded rate + external spend + amortised tools |
| 2 | Internal hours per asset, split by stage | 10.8 h (brief 1.9, SME 1.5, edit 3.5, publish 1.6, QA/promo 0.9, comms 1.4) | Toggl Track or Harvest, two weeks of honest logging |
| 3 | External spend per asset | £224 (1,400 words at £0.16) | Invoices, not estimates |
| 4 | Median calendar days, brief approved to publish | 19 days | Airtable or Notion date fields |
| 5 | Queue time as a share of calendar time | 92% | (Calendar days − touch days) ÷ calendar days |
| 6 | Revision rounds to approval | Median 2, p90 4 | Google Docs version history, or count comment threads |
| 7 | Published assets per month per FTE | 1.6 | Sitemap or CMS export ÷ FTE |
| 8 | Substantive rewrite rate | 21% of first drafts need structural rework | Editor’s tally, one field in the tracker |
| 9 | Factual and brand corrections per 1,000 words | 3.4 | Same tally, counted during edit |
| 10 | Day-90 organic sessions per asset: median, p25, p75, mean | 41 / 9 / 148 / 214 | GA4, joined to publish dates |
| 11 | Day-90 engaged session rate and median engagement time | 58%, 1m 47s | GA4 |
| 12 | Conversions per asset and cost per conversion | 143 demo requests across 68 assets, £376 per conversion | GA4 key events, or HubSpot attribution |
Numbers 10 to 12 need per-asset windows, which is why they’re the ones that get fudged. If you have the GA4 BigQuery export switched on, pull raw sessions by landing page first and worry about the windowing in the join:
SELECT
REGEXP_EXTRACT(
(SELECT value.string_value FROM UNNEST(event_params) WHERE key = 'page_location'),
r'^https?://[^/]+([^?#]*)') AS page_path,
COUNT(DISTINCT CONCAT(user_pseudo_id, '-',
(SELECT value.int_value FROM UNNEST(event_params) WHERE key = 'ga_session_id'))
) AS sessions
FROM `my-project.analytics_412903117.events_*`
WHERE _TABLE_SUFFIX BETWEEN '20251001' AND '20260930'
AND event_name = 'session_start'
AND (SELECT value.string_value FROM UNNEST(event_params)
WHERE key = 'medium') = 'organic'
GROUP BY page_path
ORDER BY sessions DESC
No BigQuery export? A GA4 free-form Explore with landing page as the dimension, session medium as a filter and a manual date range per publish cohort gets you close enough. What it produces looks like this, and the shape matters more than the absolute figures:
page_path sessions_day90
/blog/rota-software-comparison/ 6,213
/blog/holiday-pay-calculator-guide/ 1,880
/blog/tronc-schemes-explained/ 703
/blog/shift-swap-etiquette/ 27
/blog/five-rota-mistakes/ 11
...
n = 68 median 41 p25 9 p75 148 mean 214
One caveat worth building into the export now: Search Console only retains 16 months. Whatever query-level and position data you want as part of your baseline, pull it this week via the Search Console API or the official Sheets add-on and park the CSV. It will be gone before your first proper comparison.
Where AI actually lands in a £654 blog post
Now decompose that £654. Drafting, the thing AI is genuinely good at, is the freelance line: £224. Everything else is brief-writing, an SME interview, two edit rounds, chasing legal, uploading, internal linking, promo copy.
Swap the freelance draft for an AI-assisted one and you remove £224. Realistically you add editing time, because a machine draft needs more structural work than a good freelancer’s: call it two extra hours at £35, so £70. Net saving £154 per asset, or 24%. Across 68 assets a year that’s £10,472. A real, defensible, boring number, and one you can only produce because line 3 of your baseline exists.
Cycle time is where the claims get wilder and the baseline bites hardest. Touch time of 10.8 hours is about 1.5 working days inside a median 19-day calendar span. The other 92% is queue: waiting for an SME slot, waiting for legal, waiting for the designer. Cut drafting to zero and median brief-to-publish moves from 19 days to roughly 17, perhaps 18. If anyone tells you AI halved the content cycle, number 5 is how you show that what actually halved was 8% of the process. (Which, incidentally, tells you where to spend your next improvement effort, and it isn’t drafting.)
Your performance numbers will never reach statistical significance
Look again at row 10: median 41, mean 214. That gap is the whole problem. Per-piece organic performance is a heavy-tailed distribution with a coefficient of variation around 2.2 in this example. A rough power calculation, n ≈ 16 × CV² ÷ Δ², says that detecting a 25% relative lift in mean sessions needs something like 1,240 assets per arm. The team publishes 68 a year.
That isn’t an argument for giving up on measurement. It’s an argument for putting your ROI weight on cost and cycle time, where a single well-logged asset is a valid measurement and variance is low, and treating per-piece performance as a non-inferiority guardrail rather than proof of gain. Define the guardrail before you start: for example, the AI-assisted cohort’s day-90 median must stay within 15% of the baseline median, and the share of assets failing to clear the baseline p25 (9 sessions) must not rise above the baseline rate of 25%. Two breaches in a quarter triggers a review of the workflow, not a debate about whether AI “works”. For the fuller treatment of attribution windows, cohort design and control groups, our guide to measuring AI content performance and ROI goes considerably deeper than a guardrail rule.
Keep a control cohort too. Hold 25 to 30% of output on the existing human-only workflow for at least two quarters. Core updates, seasonality and your own site changes hit both arms equally; without a control you will attribute every algorithmic wobble to the new process, in whichever direction suits your prior.
Freeze it, or it rots
Snapshots decay in a specific and predictable way. A Looker Studio report set to “last 12 months” rolling will, by March, be quietly comparing your AI cohort against a baseline that already contains AI-assisted posts. Export the twelve numbers to a dated file, baseline-2026-09-30.csv, drop it in Drive somewhere boring and read-only, and write the loaded hourly rate and the date range into the file itself. Also write down what you couldn’t measure. “SME interview time not logged for 19 of 68 assets” is far more useful in nine months than a suspiciously complete dataset.
The single highest-value change to make this week is one field in the tracker: assist_level, with values none, outline, draft, full. Add it to Airtable now, before anyone needs it, and backfill none across the last twelve months. Every cohort comparison you’ll want to run for the next two years depends on that column existing from day one, and no amount of later archaeology reconstructs it. The first AI-assisted brief can go out on Thursday. It just can’t go out before the snapshot does.