5 rules I build into every multi-step AI workflow for SEO

I run several multi-step Claude SEO workflows: article refresh chains, video content pipelines, AI retrievability audits, to name a few. You might have seen some of them in one of my talks, at the AI Search & Systems Summit 2026. They all run reliably now, and produce exactly what I want them to in the way I want them to, but getting there took several rounds of running and auditing to find the blind spots, work out which step caused them, and fixing them properly.

All of what follows applies to single skills too, but I’m framing it around chains because that’s where it can get more difficult to spot where a workflow fails. If you have six steps running together, each handing output to the next, you can easily lose track of which one stopped doing its job properly.

Below is what I identified through trial and error and what I now integrate every time I build an AI pipeline for SEO.

1. Build verification into the skill itself

A pipeline of mine used to mark outdated claims with [NEEDS VERIFICATION] and stop there, because that’s what I’d set up in my overall custom rules. More than enough as a default, but inside a workflow it isn’t, since it hands the lookup back to the person the verification was meant to save time for in the first place.

So I added a web search sub-step: find the current figure first, tag it [UPDATED] with a reliable source when one turns up, and fall back to [NEEDS VERIFICATION] only when nothing reliable exists. Of course you need to define what reliable means to you in the skill file too, or the model picks for you. That way, the review flag means “I checked and couldn’t confirm“.

2. A tool that returns nothing has to say why

During a test, a transcript fetch step returned 403s in the sandbox because YouTube blocks data center IPs to prevent scraping. In the output that read as “no transcript available”, but I knew that wasn’t possible, having tested the same thing from another environment AND having manually checked. It was a false negative that would have dropped real competitor data from the analysis without me knowing.

Same thing with an SEO data tool MCP call coming back empty, which can mean no results, a bad parameter or a quota hit, or with Screaming Frog returning nothing because the MCP server wasn’t running. Three very different problems, and fixable ones, if they’re spotted.

Again, the fix was a simple instruction in the skill that forces it to report why it got nothing and stop there, rather than pass an empty result down the chain. A tool refusing, a tool not being there or a tool genuinely finding zero results are still findings.

3. Give skills real examples, so they don’t assume

The script and metadata steps in one of my pipelines were built on assumptions that seemed okay while testing, then produced output that was never quite what I wanted compared to what was already published on the channel. Shorts got treated as their own fixed structure rather than a repurposing of the long-form script, chapter names defaulted to generic labels, and I had to adjust the output every time.

The fix was a proof file with examples of gold-standard published videos, which those steps now read before they write anything.

This is the P in SCRIPT, the prompting framework I created and use for this exact reason. Situation, Character, Request, Instructions, Proof, Template. Proof is what stops the model from assuming it needs to follow a certain structure when you haven’t given it one.

4. Know where each step actually executes so you can reproduce

The same workflow ran fine in Claude Code and failed in Claude chat, because of where the skill was running. Claude Code runs locally, with access to files sitting on my machine (like when using the Google Search Console connection); Claude chat executes in a remote container. If the model has to reach something on your computer, it can’t do that in Claude chat.

Keep a record of what you build (yes, you can automate this too!), what runs where, and what each environment has access to: local files, API keys, MCP servers, and so on. Write it down and you’ll thank yourself later, instead of going back through skills, chats and sessions to retrieve something you can’t remember.

5. Push fixes right back into the skill file

This one seems obvious, but sometimes we just forget!

You run the workflow, something is wrong, you fix it, and the next time you run it the same error is back and you think “Wait, didn’t I fix this already last time?” The beauty of being human and not a machine, dare I say.

Then you have to dig through old chats to find what you did the first time. That happened more than once before I dealt with it properly.

The fix is a line in my custom instructions: when fixing a skill or pipeline output, ALWAYS ask whether the skill should be updated too. It’s already saved me so much time and rework.

If you want to go further, Claude Code’s goal-based and proactive loops can automate this instead of you deciding when to step in.

Before you chain anything, test each skill on its own

I know that it seems easier to let the AI do all the work and write the skills and save it and all, but the time you put in at the start of the process does pay off in the long run, when you know exactly what you built and how you optimized it, and that makes it easier to achieve higher-quality output and then build on top of it.

Confirm a single skill in isolation with real data, then once it works, chain it to other skills. This way you can be confident it’s handing the next step something you can trust.

Share it!