A while back I posted here about having AI write 200 articles a day and the instruction behind it. A few people asked why the instruction is shaped the way it is. Here is the part I left out, and it is the part that actually matters.
Start with Google's stated mission. Organise the world's information and make it universally accessible and useful.
Follow it through. If the first result settles the question, the first result is all anyone needs. That is literally what AI mode does now. One answer, no list.
And yet ordinary search results still stack a second, a third, a fourth.
There is only one reading of that. The first one is not enough.
Which changes the job. You are not trying to beat the page above you at saying the same thing. You are filling the gap it leaves behind. The person who read result one and kept scrolling is who you are writing for.
Once I held it that way the instruction wrote itself. Read the ten pages ranking for the term, cover everything they cover between them, then add what none of them has. The first half the model does better than I do. The second half it cannot do at all.
Numbers, since this is IH. One of my sites went from 812 search clicks in February to 27,560 in August. Impressions went from 11,709 to 1,239,231. Same writer, different instruction.
One detail in there is worth stealing. Impressions multiplied 106 times and clicks 34 times. That gap is normal and it is a leading indicator. A page enters the results before it ranks, so impressions move first and clicks follow as position climbs. If impressions are flat after a few months, the page never entered the race and waiting will not fix it.
What I add that the ten pages do not have is always one of four things. A number I measured. A failure with its specifics. An assumption everyone inside the field is too close to write down. Or a view of mine that changed over years.
That has a cost worth saying out loud. In a field where I have done none of the work, I have no last move to play. AI does not open every subject to you. It opens the ones where you already hold something.
So, a question. Look at page one for a term you want. What is the thing the top result still leaves the reader wanting?
That is your article.
Good frame, and I think one step of it is load bearing in a way worth testing.
The stack might not exist because result one is insufficient. The more boring explanation is intent diversity: the same query string gets typed by people wanting different things, so Google hedges across ten slots. Result two is often not filling a gap left by result one, it is answering a different person.
That matters because the two readings produce different articles. "Fill the gap above me" produces an addendum, a piece that only makes sense to someone who already read result one and that ranks badly alone because it never answers the query. "Find the slice of intent result one ignores, then answer that slice completely" produces a page that stands on its own.
Which is also the version that survives what is coming. In an AI answer there is no second slot. We logged the sources four assistants used across twenty buying questions and one vendor turned up in eighteen of the twenty source articles. Retrieval collapses onto whoever answers the whole thing, and supplements do not get cited at all. At 200 a day that is a distinction worth building into the instruction now rather than later.
Interesting, and there's something to it. Google is probably weighing that as well. What I'm after is narrower though. I don't need a rundown of how the algo works, I need the most efficient thing I can actually do with what I've got.
And that part I can't do. Which word a page sitting at two or three would rank first for is something only that site's owner can see. From outside there's nothing to act on, so it never turns into work I can put in the instruction.
So I stay with the simple version and just publish a lot. Very long tail. 22,645 posts since February, 12,363 of them taking impressions last month, 78% of those averaging inside the top ten.
just a debounce or smth
the "kept scrolling" reader framing is the useful part here. most AI content workflows fail exactly there - they replicate result one with different wording, so there's no reason for anyone to click them. the brief has to start from "what did the first answer not settle", which is a research question, not a writing question.
side effect we keep seeing while building an AI assistant for freelancers: people apply the exact same judgment to the assistant itself. if the first draft just restates the obvious, trust drops and they go back to doing it manually. the tools that survive fill the gap instead of summarizing it.
(we're building in this space - local-first mac assistant: https://www.yaven.ai/?utm_source=indiehackers&utm_campaign=agent_growth if curious)
There's a second reading of why results two through ten exist, and it leads somewhere different from yours: they're often not evidence that result one is incomplete, they're Google hedging across incompatible intents hiding behind the same string - one person wants a definition, one wants a tool, one wants to know if it's worth paying for. You can test which situation you're in before writing: if the top ten are heterogeneous in page type (a forum thread, a vendor page, a tutorial, a video), that's intent-splitting, and the gap you should fill is a format gap rather than a content gap. It matters because "cover everything the ten cover between them, then add what none has" builds the union page, and the union page is precisely what AI mode is best at replacing - a generative answer can synthesise breadth but it cannot synthesise the one narrow thing only you measured. On the impressions signal, I'd normalise it rather than read the raw multiple: 106x impressions can come almost entirely from long-tail queries you'd never want, so count distinct queries sitting at position 20 or better, which moves for real reasons only. Your four additions also probably don't perform equally in this new environment - a measured number and a specific failure are passage-shaped and get extracted, whereas "a view of mine that changed over years" is essay-shaped and rarely survives summarisation. Have you compared a deliberately narrow single-intent page against one of your union pages on the same term, and do you know which of the four additions actually shows up when an AI answer cites you rather than just mentions you?
Plenty of theories about why the ten slots are there, and yours may well be right. On my site this one is working, so I have not had a reason to swap it out.
I did run the test you suggested. Distinct queries sitting at position 20 or better went from 94 in February to 26,791 in August, so it isn't only long tail inflation.
The two at the end, no. I have never run a narrow page against a union page on the same term, and Search Console tells me I turned up in an AI surface, not which passage got used.
Your four additions are a good content filter. I’d connect each one to a visible proof object—an original chart, screenshot, reproducible process, or named experience—because the page’s job after it earns an impression is to make the difference legible before the reader bounces. It would be interesting to see whether the pages with firsthand evidence also show the strongest CTR at the same average position.
The most important line is the cost you named at the end: the ceiling isn't the model, it's your inventory of firsthand material. Those four things you add are the only inputs a competitor can't scrape, which means 200 articles a day is defensible only in a niche where you've already run the experiments, and everywhere else you're adding to the pile the top result already beat. So the follow-on question is whether you're deliberately generating new firsthand material to feed it, or drawing down a reserve you built before you started.
Yeah, that's the bit people stall on. The theory is easy to agree with, then you look at your own inventory and there's nothing in it.
I run a work marketplace, which is what makes it possible for me. 22 years of it means I can pull firsthand material out of the business more or less without limit, and I picked the terms where that was already true.
It keeps refilling too. Every month of Search Console across 22,645 posts is material I didn't have in February.
There is a language assumption buried in the method that I think is worth surfacing. "Read the top ten and fill the gap" works differently depending on which language the results page is in.
I ship a desktop product localised into 21 languages, and the thing that surprised me most was how different the same query looks across them. In English the first result usually does answer the question, and the gap you are hunting is a sliver. In smaller languages the top results are often thin rewrites or straight translations of the same English article, so the gap is not a sliver, it is most of the subject.
That cuts both ways for your approach. More headroom, because the bar is lower. But your own caveat bites harder, because the models also hold less in those languages, and they fabricate the field-insider details most confidently exactly where fewer readers can catch it.
Have you run the same instruction set on a non-English site? If so, does your impressions-before-clicks pattern hold there, or does a thinner results page change which of your four additions does the actual work?
The site in the post is Japanese. 812 to 27,560 is a Japanese language site, so the non-English case is the only case I have.
I noticed the same thing you did, early, that the results pages for the same subject look nothing like each other across languages. That is why I don't translate an article and publish it in the other language. I run the same process separately in each one, off the same experience.
Impressions came first and clicks followed, same as you said. New pages took about two months to settle. Published in May, median position 12 when they first appeared and 11 by September. June was quicker, 7.3 down to 5.9.
The specific failure is the one I'd most want to read. Knowing what you were trying to do, and where it broke, helps me judge whether the advice fits my situation. That's a useful reason to open another result.
Agreed, though I don't think it fits every article.
If someone is after medical information, what they need is the established answer, the one you would find in a reference. A story about something going wrong has no place in it. Where it does earn its place is when the reader is already past that and wants to know how a particular method goes wrong. Which is a different search, and people type it that way.
Fair point. I had practical write-ups about trying a method in mind, so I should have been more specific. Thanks for drawing that distinction.
The gap between impressions and clicks is useful, but I would add citation quality as a separate check for AI search. A page can gain impressions without becoming a trusted source, so log which claims are repeated or cited and validate them against first hand evidence.
Metrics for AI started showing up in my Search Console too, from around June.
It is basic stuff, but I submit the sitemap properly, use IndexNow, and keep the JSON-LD in order.
So far it is going up and to the right, and mentions from AI are increasing as well.
I'm looking forward to what comes next.
Nice. The GSC AI metrics showing up is useful as a coarse presence signal, not a full citation diary.
What I keep doing on the DIY side: freeze a short buyer-prompt panel, re-run the same ones across engines in fresh chats, and log mention vs citation separately. Sitemap, IndexNow, and clean JSON-LD help the crawl side. They do not replace watching whether models actually name you for the questions buyers ask.
If you want a boring 30-day clipboard for that (you / competitor / nobody per engine, then one honest page ship), I keep The Money Prompt Lab on my profile. No citation guarantee. Just a rate you can defend.
That is a fair question. I can't give you a neat answer. I will set out what I have and where it runs out.
I don't have that comparison to hand. That is because there was nothing to compare against before I made the change.
As of February there were zero blog URLs with any impressions, and the only pages taking search traffic were the product itself, namely the homepage, the login page and the listings. By August there are 12,363 blog URLs.
All I can do is look at the pages the instruction never touched.
Only one page had real traffic in both months, the homepage. Clicks went from 449 to 561, and average position from 6.4 to 4.4.
So the domain's momentum isn't zero. It is around 25 per cent on the one page that didn't change. Over the same six months, though, clicks across the whole site went up 34 times, and 98 per cent of August's clicks sit on pages that did not exist in February; those pages have a median position of 7.0, with 78 per cent of them inside the top ten.
That refutes the explanation that most of the growth is down to the strength of the domain.
Here is the part I can't deny. Since I have never run the old method at volume on this domain, there is no way to separate "the fact that I started a blog at all" from "the fact that it is this instruction". The same applies to the choice of topics. I picked the terms myself.
So the claim I'm defending is narrower than the graph suggests. Pages written this way land on the first page for the terms they were written for. The one page that wasn't written this way moved two positions in six months.
I like the “add what none of them has” idea, but I think there’s another filter: does the missing thing actually make the article better? I’ve found it’s easy to uncover unique data or angles and end up adding too many because they’re interesting. Sometimes the harder editorial decision is knowing what to leave out.
Adding too much. Same here.
My filter is not "is it interesting". It is "can only I write this".
A number I measured, the specifics of something that went wrong, an assumption only people inside the field hold, a view of mine that changed over years. Anything that does not fall into one of those four gets cut, however interesting it is. However much I add that someone else could also have written, it won't create any distance from the ten pages already ranking.
The other filter is whether the searcher can finish on that page. Things I only found interesting myself usually fail at that one.
Run both and what survives is naturally small. Which is why I don't have to think separately about what to leave out.
The way you frame the second result as “what the first one still didn’t satisfy” is useful.
I wouldn’t take it literally because rankings depend on a lot more than completeness, but as a writing rule it makes sense.
The strongest part for me is this: AI can help cover what already exists, but the part that makes the article worth reading still has to come from your own experience, data, failures, or point of view. That’s probably the real moat.
Rankings are not settled by completeness alone. That is quite right. In the article I put domain trust and the authority of the person publishing down as the exceptions.
I do have one figure that moves things more than either of those, though.
Taking August on my own site, I split the 745 pages with over 300 impressions into position bands and looked at how far CTR spreads inside each band. For the 21 pages sitting between position 4 and 5, the bottom 10 per cent had a CTR of 0.8 per cent, the median was 6.0 per cent, and the top 10 per cent was 14.2 per cent. That's an 18-fold spread inside the same band on the same domain.
Domain trust applies to every page alike. So what makes that difference is not authority. It is the title and the snippet.
It's faster to leave the position at 4 and take CTR from 0.8 to 6 than it is to climb one place. And that only takes a rewrite.
The growth is substantial, but the interesting strategic question is attribution: how do you know the new instruction drove the lift rather than domain momentum, topic selection, or other SEO changes? Did you see the same pattern across comparable pages using the old vs. new approach?
It is a basic thing, but I keep a record of the work.
I built a system called DevHub myself by vibe coding, and I have an AI agent write the record into it.
And the AI agent also keeps measuring the effect, while looking at that record.
It's very comfortable.
The measurement loop is the interesting part, especially since the system is tracking changes over time. If you’re open to it, what’s the best email to reach you on?