All Posts
November 25, 2025AIWorkflow

Three Frontier Models in One Month. Your Workflow Matters More

In the span of a few weeks this month, the major labs all shipped new flagship models. Each one topped the last on the benchmarks. Each launch produced a wave of posts declaring a new king.

I use these tools every single day, across writing, design, production planning, and app development. I run communications for ministry organizations, I build client work through my consulting practice, and AI sits somewhere in nearly all of it. If anyone should have been glued to the launch coverage, it is me. Here is my confession: the launches barely changed my week.

That is not a complaint about the models. The new ones are genuinely better. It is an observation about where the value actually lives, and after a month of watching smart people burn energy arguing about which lab is winning, I think the observation is worth writing down.

Switching is cheap. Workflow is not

When a better model ships, I change a setting. Ten seconds. What I cannot swap in ten seconds is everything around the model: the voice guides that keep output sounding like our brands, the prompt templates my pipelines run on, the connectors into email and community platforms and project boards, the review gates that keep a human on every send.

Let me itemize that surrounding system, because "workflow" is a vague word and the vagueness is exactly why people underinvest in it.

The voice guides came first and took the longest. Every brand I touch has a written document describing how it sounds: the vocabulary it uses, the vocabulary it refuses, how it opens, how it closes, how formal it runs, what it never jokes about. Writing those meant sitting with years of published material and getting honest about patterns I had been carrying in my head. A model without that document produces competent generic text. A model with it produces something I can nearly ship. Every model, including the three that launched this month, is transformed by the same document.

The prompt templates came next. Any task I do more than twice gets a template: the newsletter draft, the event announcement, the social cutdown of a long video, the meeting summary. Each template encodes decisions I only want to make once, like what structure the newsletter follows and what always goes in the closing block. Templates are boring. Templates are also why my output is consistent on the weeks I am tired.

Then the connectors, the plumbing that lets the assistant read my actual systems instead of my paraphrases of them. And finally the review gates, the deliberate human checkpoints before anything reaches an audience. That whole stack took months of evenings to build.

And here is the punchline: the stack took months, and it is the actual asset. It makes every model better the moment it arrives. Drop a smarter model into a mature workflow and the whole pipeline gets smarter overnight, for ten seconds of switching cost. Drop the smartest model in the world into no workflow and you get a very intelligent chat window that produces off-brand text you have to babysit. People who spent this month arguing about which lab is winning could have spent it building a system that benefits no matter who wins.

Twenty years of gear arguments prepared me for this

I have seen this movie before, just with different props. I spent two decades in video production, and video people love a spec war. Camera launches were our model launches. Every year a new body shipped with more resolution or better low light, and the forums declared the old gear obsolete, and grown adults argued about sensors the way people now argue about benchmarks.

Meanwhile, the working professionals I respected upgraded on a completely different rhythm. They asked one question: does this change what I can deliver? Sometimes it did, and they bought without sentiment. Usually it did not, and they kept shooting, because they knew where their results actually came from. Their lighting. Their audio discipline. Their shot lists. Their edit instincts. The camera was maybe fifteen percent of the outcome, and it was the fifteen percent that got one hundred percent of the arguing.

The best wedding videographer I ever knew shot on gear two generations old, and his films made people cry. His secret was not in the sensor. It was in knowing where to stand and when to press record. I never once heard a bride ask what camera he used.

Models are cameras. The workflow is the lighting, the audio, and the instinct for where to stand. The specs will keep leapfrogging forever, and the arguing will follow the specs, and the results will keep coming from the boring surrounding system that nobody posts about. If two decades behind a camera taught me anything transferable, it is to notice which fifteen percent is soaking up all the attention and go invest in the other eighty-five.

The practical takeaways

Stay model-agnostic. Keep your prompts, voice files, and context documents in a form you own, not locked inside one vendor's product.

For me this means plain files in my own storage: markdown documents for voice guides, plain text for prompt templates, ordinary folders for reference material. Nothing exotic. The test is simple: if a vendor disappeared tomorrow, or doubled their prices, or shipped a model I did not trust, could I move in an afternoon? Because I keep my materials portable, the answer is yes, and that portability quietly changes my relationship with the entire market. Every launch this month was good news for me, because any of the three new models could slot into my system. That is the position you want: a customer every lab has to keep earning, not a resident of anyone's walled garden.

Benchmark on your work, not theirs. I keep a small folder of real tasks: a newsletter to draft, a script to cut down, a data cleanup. When a new model ships, I run the folder through it. Twenty minutes tells me more than any leaderboard.

The folder deserves a little detail because it is the highest-leverage twenty minutes in my whole AI practice. It holds a handful of completed tasks from my real work, with the inputs I had and the output I eventually shipped. A newsletter with its source notes. A long transcript and the short script we cut from it. A messy spreadsheet and its cleaned version. When a new model arrives, I hand it the same inputs and compare its output to my shipped versions. The results are frequently surprising in both directions. Models that dominated the public benchmarks have stumbled on my newsletter voice. A model nobody was hyping turned out to be clearly best at cutting transcripts, which happens to be a thing I do every single week. Public benchmarks measure the average of tasks I mostly do not have. My folder measures my Tuesday.

Upgrade when it changes an outcome. Faster and smarter matters when it moves something you ship. Otherwise it is entertainment.

This month, for the record, my folder test did move me on exactly one thing: one of the new models is meaningfully better at long-document work, which improves a real weekly task of mine, so it earned a place in that pipeline. The other launches were, for my purposes, entertainment. I say that with no disrespect. I enjoy the entertainment. I just do not confuse it with progress on my actual work, and the distinction is the whole discipline.

The question I ask instead of "which model is best"

When someone asks me which model they should be using, I have started answering with a different question: what does your system look like around whatever model you choose? Because in every organization I consult with, that is where the gap is. Nobody is losing because they picked the second-best model. The differences between frontier models this month are real but marginal. People are losing because they have no voice documentation, no templates, no connected tools, and no review process, which means they are getting maybe a tenth of the value out of whichever model they picked.

The uncomfortable part is that workflow building feels like homework while model news feels like sports. One produces compounding returns and no dopamine. The other produces dopamine and no returns. I understand the pull. I read the launch threads too. But if I audit where my actual productivity gains came from over the past year, close to none of it traces to model upgrades, and nearly all of it traces to evenings spent writing voice guides and wiring connectors. The boring investment won by a landslide.

The model race is real and I am glad it is happening. Competition is why capability keeps jumping, and every jump makes my system more valuable. But if you are responsible for output, not opinions, the leaderboard is not your scoreboard. Your calendar is. Look at what you shipped this month versus last month. If that number is not moving, no launch announcement is going to move it for you. The system will. Go build the system.