Engineering for the ChatGPT and Perplexity citation layer.
A 12-month structured experimentation programme on what makes content cite-able by ChatGPT, Perplexity, Gemini, and the next generation of AI retrieval systems — the chunking models, entity signals, and source-authority patterns that now shape every engagement.
Why I started running these
In mid-2024, the discoverability landscape changed faster than the SEO playbook could update. ChatGPT, Perplexity, and Gemini were directing meaningful share of high-intent queries — and the SEO community’s response was mostly anecdote and speculation. Practitioners had strong opinions; almost nobody had structured experimental data.
I started this programme because the work I was doing for clients was accumulating questions I couldn’t answer rigorously. Why did this page get cited and the equivalent competitor page didn’t? What changed when we restructured the H2s? How did entity signals actually affect retrieval? The only honest way to answer was to run experiments.
The experimental framework
Building useful experiments in AI search is structurally hard. The retrieval systems are black boxes. Sample sizes are small because queries are expensive to track. Outputs are noisy at any reasonable scale. Most of what gets written about “what works” in AI search is selection bias dressed as insight.
The framework I built was designed to produce signal despite these constraints:
- Paired pages, not isolated tests. Every optimised page has a structurally similar control page on the same site or sibling site. Single-page results are meaningless; paired comparisons surface signal.
- 340 fixed-query tracking set. A pre-registered query set, sampled monthly across all 5 models. Same queries every month, so trend data is comparable.
- Pre-committed hypotheses. Before each cohort of pages launches, I write down what I expect to happen and why. Hindsight rationalisation is the biggest risk in experimentation; pre-commitment is the cheapest defence.
- Sufficient duration. Each experiment runs 90+ days. Citation surface shifts slowly — anything faster is usually noise.
“Most of what’s published about AI search is selection bias. The interesting work is in pre-registered experiments with control pages — and that takes patience the field hasn’t built yet.”
The three layers of citation engineering
The experimental work surfaced three layers of structural intervention that consistently affected citation rates. Each layer is independently testable and stacks with the others.
What surprised me
Two findings ran counter to what I expected going in:
Recency mattered less than I assumed. For evergreen queries, content age was nearly irrelevant. The retrieval systems weighted structural quality much more heavily than publication date for non-time-sensitive topics.
Domain authority traditional metrics mattered less than functional authority signals. A small site with original benchmarks consistently outperformed large sites with editorial content on the same query. The lesson: AI retrieval rewards a different kind of authority than traditional SEO trained us to optimise for.
How the experiments actually run
Each experiment cycle takes 90–120 days. The mechanics:
- Cohort selection. 6–10 pages selected for a given hypothesis. Half optimised, half held as control. Pages chosen to be as structurally similar as possible — same topic depth, same domain authority bracket, similar baseline traffic.
- Pre-launch baseline. 30 days of citation tracking before any changes. Establishes the noise floor for each page.
- Optimisation deployment. The intervention ships on the test set. Control pages stay untouched. Documentation captures what changed and why.
- Tracking window. 90 days of monthly citation snapshots across 5 models. The first 30 days are usually noisy; weeks 6–12 are where signal stabilises.
- Pre-registered analysis. Compare cohorts against pre-committed hypotheses. Reject hypotheses that don’t clear the noise floor, even if they’d be convenient to confirm.
This is much slower than the “test things and tweet conclusions” pattern that dominates AI search content. It’s also the only way to produce findings I trust enough to deploy in client engagements.
What the data shows
The headline finding is the one I’d most like to caveat: structural optimisation produces median 4.2× citation-rate lift across optimised cohorts, with significant variation by query type and model.
The variation matters. For high-intent commercial queries (where Perplexity and ChatGPT are most active), the lift was closer to 6×. For exploratory informational queries, closer to 2.5×. For navigational queries with strong brand entities already established, optimisation moved the needle barely at all.
- 4.2× median cite-rate lift across optimised cohorts — varied substantially by query type and model, but consistently positive across every experiment that completed its 90-day window.
- 6 reproducible citation patterns identified — each documented with the experimental conditions, the intervention, and the observed lift. All now in active use in client engagements.
- 340-query tracking set running continuously, providing the dataset for future hypotheses and a baseline that’s now 12+ months deep.
- Operating insight that drives engagement work. Every client engagement now starts with a citation-surface audit using this framework. The experimental programme became the source material for productised consulting work.
What I’ve learned about running personal labs
Three reflections from running this for 12 months:
- The discipline of pre-registration is more important than the experiments themselves. Half the value of this programme isn’t the findings — it’s the habit of writing down predictions before running tests. The number of times I’d have rationalised noise into “insight” if I hadn’t pre-committed is genuinely uncomfortable to admit.
- Personal labs compound differently than client work. Client engagements give you depth on a single system. A personal lab gives you breadth across patterns that surface only when you can run experiments without commercial pressure. The two are complementary; neither replaces the other.
- Most of what gets shared on AI search is wrong, but not maliciously. It’s wrong because the field is moving faster than rigorous experimentation can keep up. The honest thing is to keep running experiments, keep publishing what you find, and keep being explicit about confidence levels.
This work is ongoing. New cohorts ship each quarter. The findings will continue to evolve as the retrieval models do. The framework is more durable than the specific results.