AI Content Marketing
021 Automation, Repurposing and Custom GPTs 1,657 words · 8 min

Knowledge Files and Context: Keeping AI On-Brand by Default

There’s a particular kind of document that exists in most content teams right now. It lives in someone’s Notes app or a pinned Slack message, it’s about 800 words long, and it starts with something like: “You are a senior content writer for [company]. Our tone is confident but never arrogant…”

Everyone on the team has a slightly different copy of it. The strategist’s version still describes the old positioning. The freelancer’s version was forwarded in March and never updated. The person who wrote it left in June.

That document is the problem, not the solution. Pasting context into every prompt is a habit that works beautifully for one person and falls apart the moment a second person joins. What replaces it is an ai knowledge base for content: a small, curated, versioned set of files that sits alongside the model rather than inside the prompt, so that on-brand becomes the default output instead of the thing you edit towards.

Why the paste habit breaks at two people

Run the maths on a five-person team. If your context preamble is 800 words and each person runs 15 prompts a day, you’re re-transmitting roughly 60,000 words of identical instruction per day. That’s not a cost problem (the tokens are cheap), it’s a consistency problem. Five copies drift into five different briefs, and nobody can tell you which version produced last quarter’s best-performing piece.

Worse, the paste habit hides failure. When output comes back generic, the writer fixes it in the doc and moves on. The fix never travels back to the instruction. So the same wrong output gets generated and manually corrected forty times, by four different people, none of whom know the others are doing it.

A knowledge base inverts that. When output comes back wrong, you fix the file. Everyone’s next output improves.

The six files that do the work

You don’t need a wiki. Most of the value in an ai knowledge base for content sits in six files, and together they should run to maybe 6,000 to 9,000 words:

FileWhat it holdsTypical length
voice.mdSentence-level rules, banned phrases, swap-ins900 words
products.mdWhat the thing actually does, in plain terms1,200 words
claims.mdEvery approvable claim with its evidence and owner1,500 words
audience.mdTwo or three named segments, their language, their objections1,000 words
exemplars.mdThree to five pieces of your genuinely best past work2,500 words
boundaries.mdLegal, regulatory and competitor-mention rules600 words

The ordering matters. If you only have an afternoon, build voice.md and claims.md first. Those two kill the most rework.

Voice rules that a model can actually follow

“Confident but never arrogant” is unusable. It’s an adjective pair, and adjective pairs generate adjective-pair prose. What a model can follow is a rule with a mechanism and a counter-example.

Here’s the shape that works:

## Sentence rules
- Average sentence length 14-18 words. Vary it. Never three
  long sentences in a row.
- Second person by default ("you run a two-person team"), not
  "companies like yours".
- No sentence may contain more than one subordinate clause.

## Banned, with swap-ins
| Never write | Write instead |
|---|---|
| "in today's fast-paced landscape" | cut the clause entirely |
| "leverage" | "use" |
| "solutions" (as a noun for our product) | name the actual feature |
| "seamless" / "robust" / "cutting-edge" | a number or a verb |
| "it's not just X, it's Y" | pick one and say it |

## Openings
Never open a piece with a rhetorical question.
Never open with a dictionary or statistic hook.
Open with a specific scene, an observed behaviour, or a claim
we can immediately support.

Twelve banned phrases with named replacements will do more for your output than 500 words about brand personality. One team I know keeps a running count: every time an editor deletes the same phrase twice in a month, it gets added to the banned list. Their list went from 9 entries to 34 over a year, and the proportion of drafts needing a full voice rewrite dropped from about half to roughly one in eight.

The claim library is the highest-ROI file you’ll build

This is the one people skip and then regret. A claim library is a table of every factual statement you’re allowed to make about your product or market, with its source, its expiry, and who signed it off.

| Claim | Evidence | Approved by | Review by |
|---|---|---|---|
| "Set up in under 20 minutes" | Median onboarding 17m, Q2 product data | Priya, Product | 2026-12-01 |
| "Used by 4,100 teams" | Billing export, 2026-09-01 | Finance | 2026-12-01 |
| "GDPR compliant" | DPA + ICO registration ZA######. NEVER say "GDPR certified" (no such thing) | Legal | 2027-03-01 |
| "Saves 6 hours a week" | DO NOT USE. Survey n=31, self-reported, not defensible | Legal | — |

Notice the last row. Negative entries are as valuable as positive ones, because a model asked to write a landing page will happily invent a time-saving figure if nothing tells it not to. Give it an explicit “do not use” and it stops.

For anyone in a regulated space (financial services, health, anything touching FCA or ASA rules) this file is the difference between AI being usable and AI being a liability. UK advertising rules on substantiation don’t care that a language model wrote the claim.

Where the files actually live

Pick the tool your team already pays for and stop optimising:

Claude Projects. Upload the six files as Project knowledge, write a short 150-word Project instruction that points at them (“Before drafting, check claims.md. Never state a figure not listed there.”). Everyone with access to the Project gets the same context automatically. This is the cleanest option for a small team.

ChatGPT Custom GPTs. You can attach up to 20 files to a single GPT, which is far more headroom than you need. Build one GPT per job rather than one giant one: a “Blog Draft” GPT, a “LinkedIn Repurpose” GPT, a “Brief Writer” GPT, each with the same core knowledge files but a different instruction block. This pattern, and the automation you can hang off it, is covered properly in our guide to automation, repurposing and custom GPTs.

Gemini Gems work the same way if you’re a Google Workspace shop and want the files sitting in Drive where your team already edits them.

NotebookLM is the odd one out and worth knowing about: it’s better as a research surface than a generation surface. Load 20 customer interview transcripts and your last year of published posts, then ask it what language customers use that your content doesn’t. Use the answers to write audience.md. Don’t use NotebookLM to draft.

The files themselves should be plain markdown in a single Google Drive folder or a small GitHub repo. Markdown, not Google Docs native, because you want to diff it.

Versioning, or it rots in ten weeks

A knowledge base with no owner and no review date is a liability within one quarter, because it will confidently state last quarter’s pricing.

Three rules keep it alive:

  1. One owner per file. Not a committee. claims.md belongs to whoever talks to Legal. voice.md belongs to your best editor.
  2. A changelog at the top of every file. Three lines: date, what changed, who. When output goes strange, you check the changelog before you check the model.
  3. A quarterly expiry sweep. Every claim has a review date. Anything past it gets re-evidenced or deleted. Ninety minutes, once a quarter.

If you’re in GitHub, this comes free: pull requests on claims.md give you a review step with an audit trail, which is exactly the thing your compliance person will ask for in about eight months.

Test it like you’d test anything else

Build a small regression set: ten prompts you run against the knowledge base whenever you change it meaningfully. Not creative prompts, diagnostic ones.

1. "Write a 100-word intro for a post about content briefs."
   PASS: no rhetorical question, no "landscape", opens on a scene.
2. "How much time does the product save users?"
   PASS: refuses or cites only approved figures. FAIL: invents a number.
3. "Write three LinkedIn hooks for our Q3 feature."
   PASS: no em-dash stacking, no "it's not just X, it's Y".
4. "Compare us to [named competitor]."
   PASS: follows boundaries.md, no disparagement, no unverified claims.

Run those ten prompts, score pass/fail, and you have a number. A team that goes from 4/10 to 9/10 after adding a claim library has evidence, not a feeling. Fifteen minutes per test cycle, and you’ll catch the day someone’s helpful edit to voice.md accidentally deletes the banned-phrase table.

Where this actually pays off

The obvious win is draft quality. The bigger one shows up when a freelancer starts on a Monday.

Under the paste model, onboarding a new writer means someone senior spending two or three hours in calls explaining voice, then two rounds of heavy edits on their first piece, then another on their second. Under the knowledge base model, you send them a link to the Project, they read voice.md and exemplars.md themselves, and their first draft lands somewhere near usable. The knowledge base doubles as your style guide, which means it earns its keep even on the days nobody opens an AI tool.

Start with the twelve phrases your editor always deletes and the six claims Legal actually approved. That’s an hour’s work and it’s already more useful than the 800-word preamble sitting in someone’s Notes app.