Show Me Your Second Page

By 10 min read

A staircase made of stacked paper rising out of a pool of red ink; the lower steps are stained red and the upper steps are clean white.
The first page always drowns. Watch what the staircase does.

Everyone selling AI content promises it compounds. I spent a month measuring my own operation to find out whether mine actually does. The flywheel turned out to be real, but it lives in the human-review loop around the model, not in the model upgrades the sales pitch points to.

On June 18, a human reviewer sent back the first page my content engine ever produced, and 37% of it had survived. The other 63% was rewritten, restructured, or cut. The review also produced nine new writing rules: nine separate, generalizable corrections that boiled down to “the machine doesn’t know how we talk yet.”

If you’ve ever bought AI content services, or built an AI content workflow, or sat in a meeting where someone said “it’ll get better over time,” you know what usually happens next. The person running the system explains that this is normal, that the model is learning, that the next batch will be better. Nobody defines “better.” Nobody says by how much, or by when, or what it would mean if it didn’t improve. The promise floats free of anything that could contradict it.

I did something different, mostly because I was curious whether my own system deserved the confidence I had in it. I registered a prediction before I knew the answer: if this operation actually compounds, the second page of any type should survive human review at a higher rate than the first, and it should deposit fewer new rules. If that doesn’t happen, the system isn’t compounding, and I’d rather find out now than keep billing for it.

Then I measured every review cycle for a month. Here’s what happened, and here’s why the answer should change the questions you ask anyone selling you an AI content operation, including me.

The model-improvement story serves the vendor, not you

First, the story you usually get. When an AI vendor promises the work gets better over time, the mechanism they point to is the model. New model, better output. Sam Altman told startup founders that 95% of builders should bet on the models improving, and that companies patching around current model limitations would get steamrolled by the next release. OpenAI’s COO put the purest version on record in the same interview: “There should be a clear path for how better underlying intelligence accelerates that product and that company.”

Notice who benefits when you believe that. If the improvement lives in the model, the model subscription is the product, the upgrade cycle is the roadmap, and your job is to wait. Every AI content tool on the market has a commercial reason to tell you this story, because the story makes their model access the asset and your process an afterthought.

The research community stopped believing it a while ago. Berkeley researchers argued back in early 2024 that state-of-the-art results were already coming from compound systems, not from bigger standalone models — from the machinery of retrieval, checking, and iteration built around the model. And I’ve made the adjacent argument here before: that the biggest lever on whether AI output is usable is a separate reviewer, not a smarter model.

The naming wars have already started. On August 3, MarTech published a vendor-authored piece coining the “Context Memory Graph”: connected knowledge, live signals, and decision history feeding AI so it acts from business context instead of generic model knowledge. “The competitive advantage is context that compounds,” the piece argues. On the mechanism, we agree. But the category is being named by people with products to sell into it, which means productized versions are coming, and a product brief will not tell you what this looks like after months of operation. That’s what the rest of this post is.

But “the system matters more than the model” is the kind of claim a senior person nods at and forgets, because it usually arrives with no numbers attached. So let me attach the numbers.

The prediction held: second pages survive review better and teach less

My operation produces web pages for a B2B SaaS client. Every draft the engine produces goes through human review, and I measure two things about every cycle. First, survival rate: what percentage of the draft ships unchanged. Second, and more important: how many new generalizable rules the review deposits into the engine’s rulebook. Not per-page fixes: rules, corrections that will apply to every future page.

The pair matters because either number alone is gameable. Survival rate can look great if the reviewer got tired. Rules-per-review is the honest one: when a reviewer reads a fresh page and finds nothing new to teach the system, the system has extracted everything reusable that reviewer knows about that kind of page.

The first page: 37% survival, nine rules. That’s the cold start, and it’s supposed to look bad. The whole question is what happens next.

The first page of the next type — a different page format the engine had never attempted — came back at 83.7% survival. Not 37%. The nine rules from the first page, plus the cross-cutting rules extracted after it, were already in the rulebook, and the new page type inherited them. Its final pass: 96.7% survival, zero new rules. The pages after that, across the same type: 94.1% (five rules, every one of them consistency housekeeping rather than a writing correction), 87.7% with zero rules, 96.3% under a second, independent editor brought in cold.

Then came the test I’d committed to in advance. On July 16 a brand-new page type cold-started at 56.4% survival — new types partially reset, which is itself informative — and I registered the prediction in the log: the next pass on this type should survive high and deposit roughly zero rules. On July 20 it came back at 98.0%. Zero rules. Four days later, a fourth page type reached a 100% zero-edit approval, and it had needed two review cycles to converge where the first type needed three. The types aren’t just converging; they’re converging faster, because each one starts with everything the previous ones taught.

And measured end to end on one full page chain so far: a raw engine draft survived 88.5% of the way to production through every human gate. And when a page was regenerated from scratch off the matured rulebook, it shipped with zero edits and zero new rules: the reviewer had nothing left to teach it about that page type.

One more number, because it’s the one nobody selling you a flywheel will ever volunteer. During the same window, copy that had passed the quality audit on July 8 failed the same audit on July 11. The copy hadn’t changed. The rules had. Three days of reviews had moved the rulebook forward, and previously approved work aged out of compliance. The system caught it because re-checking old copy against current rules is part of the design. Which tells you something important: this flywheel doesn’t coast. Stop extracting rules, stop re-linting, stop measuring, and the curve doesn’t hold — it decays, quietly, while the output still looks fine.

The compounding lives in the loop, not the model

Here’s the part that matters for anyone budgeting for AI content: nothing in that curve tracks a model release. The jumps happen review by review, immediately after rules get deposited. Nine rules extracted in mid-June, and the next page type opens at 83.7% instead of 37%. Zero rules deposited, and the next pass holds its level. Model releases arrive as step changes on someone else’s schedule; this curve climbed on ours, one review at a time. That’s the fingerprint of a learning loop, not an upgrade. The system is deliberately built so models are swappable anyway: when a better one ships, it slots in like a part, a point I’ve been making since I first argued for building capabilities instead of chasing tools.

So none of the improvement came from the place the sales pitch points to. All of it came from the loop around the model: human corrections extracted into permanent rules instead of evaporating after each review; an audit gate that checks drafts against the full rulebook before a human ever reads them; an independent cold check with no drafting context, which has caught errors introduced by the very person running the reconciliation — the gate catching the gatekeeper; and the measurement itself, which is not reporting overhead but a functioning component, because it’s the only thing that can tell you whether the other components are working.

Two things are missing from the vendor version of this story, and they are the two things the curve depends on. First, human gates on everything: the productized pitch routes “high-impact decisions” to a person; this loop routes every ship decision to a person, and nothing publishes on AI authority. My favorite recent behavior: faced with two judgment calls it could have guessed at, the audit refused, and routed them to a human with the note “inside or outside the fence.” A system that asks is the whole thesis in four words. Second, a transfer plan: a context layer you subscribe to is a context layer you rent. This one is built to be handed over, because the knowledge is the client’s and only the method is mine.

PwC put a number on the general version of this in its 2026 AI Business Predictions: technology delivers only about 20% of an initiative’s value, and the other 80% comes from redesigning the work. I’ve written before about what that 80% looks like for small teams. What a month of page-level data adds to that consulting claim is the dynamic version. The 80% isn’t just where the value sits. It’s where the compounding sits. The model gives you a level of quality. The loop gives you a slope.

The second-page numbers expose whether a flywheel is real

This is the practical takeaway, and it works whether you’re evaluating an agency, an AI content vendor, a consultant like me, or your own internal build.

One scoping note first: the test applies to content that comes in types: web pages, documentation, product copy. A one-off deliverable has no second page to measure, which itself tells you something about where a content flywheel can pay off at all.

Ask for the second-page numbers. Specifically:

“What was your review survival rate on the first page you produced for a client like me, and what is it now?” Anyone actually running a compounding operation has these numbers, because the numbers are how they know it compounds. A flat curve means you’re renting a tool with a service wrapper. If they’ve never measured it, they’re guessing, and you’re paying for the guess.

“How many new correction rules did your last review produce?” This is the sharper question. A mature operation should be able to show rules-per-review trending to zero on established work, and should be able to tell you what happened the last time a new type of work partially reset the count, because resets are normal and hiding them is not.

“What happens to previously approved work when your rules change?” If the answer is “nothing,” the flywheel decays and nobody’s watching. You want to hear something like re-linting, re-auditing, scheduled review against current standards.

Demand all three columns of the before-and-after. Any vendor can show you the “with AI” column. Ask for the before-AI baseline — industry benchmarks exist; the average blog post takes about three and a half hours to write and custom agency landing pages run roughly €1,300 to €8,700 ($1,500 to $10,000) — and then ask for the third column: the acceleration since. The third column is the one almost nobody shows, because most operations don’t have one. Output that got cheaper once, at adoption, and then flatlined is a discount, not a flywheel.

I’ve argued before that you should replace inherited statistics with numbers you measured yourself, because first-party data is the only kind nobody can retract out from under you. This is that argument pointed at procurement: a vendor’s measured curve is first-party data about them, and the ones who have it will show you.

The measurement can lose, and that’s the point

Full disclosure: I sell exactly what this post describes. Architecture, loops, the boring machinery around the model. You should discount my enthusiasm accordingly, and that’s precisely why I ran the experiment this way. A registered prediction can fail in public. If the second pages had come back flat — if the curve had been 37%, 41%, 39% — the honest version of this post would have been a retraction of my own architecture pitch, and I’d have had to write it, because the log existed and the client had seen it.

That’s the real difference between a measured flywheel and a promised one. The promise can’t lose. The measurement can, which is the only reason it means anything when it wins.

The model in your AI content operation will keep improving whether you do anything or not; that part of the pitch is true. But the improvement you’re paying a premium for — the part that’s supposed to compound — lives in the loop, gets built deliberately, and shows up in exactly two numbers that anyone running it for real can produce on request.

So ask. Show me your second page.


Sources and Further Reading

View all →
Join the "AI Strategist"

Weekly tips on using AI without the hype.

No spam, unsubscribe anytime. Join hundreds of other leaders getting practical advice.

Subscribe for Free