All Posts
February 24, 2026AIWorkflow

Prompt Systems, Not Prompts

Somewhere in your organization, someone has a magic prompt. It lives in their notes app, it produces great results, and when they leave, it leaves with them.

That is the state of AI in most small organizations I walk into: individual cleverness instead of shared capability. One person has quietly gotten good at this, everyone else pastes and prays, and the organization's actual competence with these tools is whatever that one person remembers on a given Tuesday. I know this pattern intimately because I was that person. My notes app was a junk drawer of prompts that worked once, half-remembered variations, and clever phrasings I could never quite reconstruct. The results were good. The system was a coin toss.

The fix is to stop thinking in prompts and start thinking in prompt systems.

How I got dragged into this

The forcing function for me was managing communications for more than one brand at a time. A ministry organization with a warm, pastoral voice. A consulting practice with a direct, practical one. Community platforms, email lists, social channels, each with its own conventions. When it was just me and one brand, the magic-prompt approach limped along. The moment I was producing for several voices in the same week, it collapsed. I would catch the consulting brand sounding suspiciously pastoral, or a ministry email opening with a marketer's hook, and I would realize the model was not the problem. My inputs were the problem. Nothing was written down. Every session, I was reconstructing the brand from memory and hoping.

The week I finally wrote the first voice document, the quality change was immediate and a little embarrassing. Everything I had been holding in my head, poorly, was suddenly on paper, consistently. That document became the seed of the system I now build for every organization I work with.

What a prompt system looks like

Mine has three layers, and none of them is complicated. That is a feature. If a layer requires a specialist to maintain, it will die the month that specialist gets busy.

Context documents. A voice file for each brand: how we sound, what we never say, real examples of our best writing. The examples matter more than the adjectives. Every organization describes itself as authentic and warm, which tells a model nothing. Five real paragraphs of your best past writing tell it almost everything. The never-say list is equally load-bearing: the terms we avoid, the phrasings that are off-limits, the claims we do not make. Then a facts file: names, dates, titles, terminology, the stuff models get subtly wrong. Which spellings are correct. What the departments are actually called. Who holds which role this year. Subtle wrongness is the most dangerous failure mode in AI drafting, because it sails right past a skim. The facts file is how you catch it before it ships.

These are documents I own, stored where the whole team can reach them, fed into any model I happen to be using. That last clause is strategy, not convenience. The context lives in files, not inside any one tool's settings, which means when a better model shows up, and one always shows up, the entire system moves with me in an afternoon.

Task templates. For each recurring job, a written prompt that references the context documents and specifies the output shape. Draft the weekly email. Turn this transcript into three social posts. Summarize this meeting into decisions and action items. Each one is a file, not a memory.

The phrase specify the output shape is where most homegrown prompts fall down, so let me make it concrete. A weak template says "write some social posts from this transcript." Mine says: three posts, each under a stated length, each anchored to one specific idea from the source rather than a general summary, no more than one question ending, written against the voice file, delivered as a draft for review. The template also says what to do when the input is thin: if the transcript does not contain three strong ideas, return two posts and say so, rather than padding. Telling the model how to fall short honestly is one of the highest-value sentences you can put in a template, and almost nobody writes it.

A review gate. Every template ends the same way: the output is a draft, a human approves anything public. Written into the process, not left to good intentions. Good intentions are what erode on a busy Thursday. The gate holds because it is structural: drafts land in a place that is not the publish button, and a named human moves them across. Which human is written down too. A review step without a name attached is a review step that will quietly stop happening.

Treat them like assets, because they are

Version them. When a template produces a weak draft, the reflex to build is fixing the template, not just the draft. This is the single habit that separates a prompt system from a prompt pile. Fixing the draft is spending. Fixing the template is investing. The draft is gone tomorrow. The template improvement pays you every week from now on. When I notice myself making the same manual edit twice, that edit becomes a line in the template that same day, and that class of correction is gone forever.

Test them when models change. Model updates shift behavior in quiet ways, and a template tuned for one model's habits can drift on another. My check is lightweight: each template folder keeps an example input with a known-good output, and when anything changes underneath, I run the pairs through and compare. Twenty minutes, a few times a year, and upgrades become boring, which is what an upgrade should be.

Share them. The whole point is that the newest team member gets the same quality output as the person who wrote the prompt. This is the part I care most about after years of building volunteer teams in church production, where the entire discipline was making excellence survivable by ordinary people. A brilliant operator who keeps it all in their head is a single point of failure with a nice personality. The prompt system is the same move I used to make with camera cards and laminated runbooks: take what the best person knows, write it down, and suddenly the team's floor rises to the level of its documentation.

What goes wrong, and how to avoid it

Having now set these systems up for several organizations, I can report the failure modes are consistent.

The first is building the cathedral. Someone gets inspired, spends two weeks producing forty templates, and burns out before any of them are tested against real work. Start with three: your most frequent writing task, your most annoying repetitive task, and your riskiest one. Run those for a month. Let the system earn its next template.

The second is the stale facts file. Facts rot. People change roles, dates pass, programs get renamed, and a facts file nobody owns becomes a liability that confidently injects last year into this year. The fix is boring and works: one named owner, and a recurring fifteen-minute review on the calendar, monthly or quarterly depending on how fast your world moves.

The third is the silent workaround. A team member finds a template producing weak results, and instead of flagging it, they quietly go back to their own freelance prompting. The system looks healthy and is actually being routed around. You prevent this culturally, not technically: make it explicitly cheap to say a template is not working, and celebrate template fixes the way you would celebrate any process improvement. The complaint is a gift. It means someone cared enough to notice.

The compound effect

After a year of working this way, the compound effect is obvious, and it shows up in places I did not predict.

New tools slot into the system in minutes because the system was never about the tool. When I want to try a new model or a new assistant, I point it at the same context documents and templates and evaluate it on the same example inputs. The trial takes an afternoon instead of a month, because I am not rebuilding my working knowledge inside each new product. I am just plugging infrastructure into a new outlet.

Team members produce on-brand work from week one. I have watched this happen with someone who had never touched an AI tool professionally: template, context documents, review gate, and their first week's drafts needed the same light-touch review as anyone else's. The expertise was in the system. The person supplied judgment and care, which is exactly what humans should be supplying.

And nothing important lives in one person's notes app. Including mine. If I stepped away tomorrow, the voice files, the templates, the review process, the example pairs, all of it stays, documented and runnable. I spent two decades in production learning that the real test of any system is whether it survives your vacation. The prompt system passes. The magic prompt never did.

One more effect, quieter than the rest: writing the system made us better at the work itself, before any model touched it. You cannot write a voice file without deciding, at last, what your voice actually is. You cannot define good output for a template without confronting how vague your standards were. Half the value of the exercise turned out to be organizational self-knowledge that had never been forced onto paper.

A prompt is a trick. A prompt system is infrastructure. Build the second one.