Ongoing · personal lab Experiment 2025 — Present

Engineering for the ChatGPT and Perplexity citation layer.

A 12-month structured experimentation programme on what makes content cite-able by ChatGPT, Perplexity, Gemini, and the next generation of AI retrieval systems — the chunking models, entity signals, and source-authority patterns that now shape every engagement.

340
Queries tracked
Across high-intent SaaS categories
4.2×
Cite-rate lift
Median across optimised vs control pages
5
Retrieval models
ChatGPT, Perplexity, Gemini, Claude, You.com
6
Citation patterns identified
Now used in client engagements
01 · Context

Why I started running these

In mid-2024, the discoverability landscape changed faster than the SEO playbook could update. ChatGPT, Perplexity, and Gemini were directing meaningful share of high-intent queries — and the SEO community’s response was mostly anecdote and speculation. Practitioners had strong opinions; almost nobody had structured experimental data.

I started this programme because the work I was doing for clients was accumulating questions I couldn’t answer rigorously. Why did this page get cited and the equivalent competitor page didn’t? What changed when we restructured the H2s? How did entity signals actually affect retrieval? The only honest way to answer was to run experiments.

02 · Strategy

The experimental framework

Building useful experiments in AI search is structurally hard. The retrieval systems are black boxes. Sample sizes are small because queries are expensive to track. Outputs are noisy at any reasonable scale. Most of what gets written about “what works” in AI search is selection bias dressed as insight.

The framework I built was designed to produce signal despite these constraints:

  • Paired pages, not isolated tests. Every optimised page has a structurally similar control page on the same site or sibling site. Single-page results are meaningless; paired comparisons surface signal.
  • 340 fixed-query tracking set. A pre-registered query set, sampled monthly across all 5 models. Same queries every month, so trend data is comparable.
  • Pre-committed hypotheses. Before each cohort of pages launches, I write down what I expect to happen and why. Hindsight rationalisation is the biggest risk in experimentation; pre-commitment is the cheapest defence.
  • Sufficient duration. Each experiment runs 90+ days. Citation surface shifts slowly — anything faster is usually noise.

“Most of what’s published about AI search is selection bias. The interesting work is in pre-registered experiments with control pages — and that takes patience the field hasn’t built yet.”

Y Yash Tulsyani · Personal Lab
03 · Architecture

The three layers of citation engineering

The experimental work surfaced three layers of structural intervention that consistently affected citation rates. Each layer is independently testable and stacks with the others.

System architecture
Layer 01
Chunkability
How retrieval systems segment your content during indexing. Self-contained sections with clear headings, standalone claim sentences, and explicit list structures consistently outperform flowing prose at chunk-extraction time.
Layer 02
Entity signals
How clearly your content disambiguates the entities it discusses. Schema.org markup, consistent canonical naming, and sameAs links to authoritative sources move citation rates measurably.
Layer 03
Source authority
How retrieval systems weight your domain as a credible citation. Original data, functional demonstrations, and named-author bylines all moved citation rates more than backlink counts.

What surprised me

Two findings ran counter to what I expected going in:

Recency mattered less than I assumed. For evergreen queries, content age was nearly irrelevant. The retrieval systems weighted structural quality much more heavily than publication date for non-time-sensitive topics.

Domain authority traditional metrics mattered less than functional authority signals. A small site with original benchmarks consistently outperformed large sites with editorial content on the same query. The lesson: AI retrieval rewards a different kind of authority than traditional SEO trained us to optimise for.

04 · Execution

How the experiments actually run

Each experiment cycle takes 90–120 days. The mechanics:

  • Cohort selection. 6–10 pages selected for a given hypothesis. Half optimised, half held as control. Pages chosen to be as structurally similar as possible — same topic depth, same domain authority bracket, similar baseline traffic.
  • Pre-launch baseline. 30 days of citation tracking before any changes. Establishes the noise floor for each page.
  • Optimisation deployment. The intervention ships on the test set. Control pages stay untouched. Documentation captures what changed and why.
  • Tracking window. 90 days of monthly citation snapshots across 5 models. The first 30 days are usually noisy; weeks 6–12 are where signal stabilises.
  • Pre-registered analysis. Compare cohorts against pre-committed hypotheses. Reject hypotheses that don’t clear the noise floor, even if they’d be convenient to confirm.

This is much slower than the “test things and tweet conclusions” pattern that dominates AI search content. It’s also the only way to produce findings I trust enough to deploy in client engagements.

05 · Results

What the data shows

The headline finding is the one I’d most like to caveat: structural optimisation produces median 4.2× citation-rate lift across optimised cohorts, with significant variation by query type and model.

The variation matters. For high-intent commercial queries (where Perplexity and ChatGPT are most active), the lift was closer to 6×. For exploratory informational queries, closer to 2.5×. For navigational queries with strong brand entities already established, optimisation moved the needle barely at all.

Median cite-rate · 12 months across cohorts Google Search Console · raw data
Month 1 Month 6 Month 12 Month 18 Month 24
  • 4.2× median cite-rate lift across optimised cohorts — varied substantially by query type and model, but consistently positive across every experiment that completed its 90-day window.
  • 6 reproducible citation patterns identified — each documented with the experimental conditions, the intervention, and the observed lift. All now in active use in client engagements.
  • 340-query tracking set running continuously, providing the dataset for future hypotheses and a baseline that’s now 12+ months deep.
  • Operating insight that drives engagement work. Every client engagement now starts with a citation-surface audit using this framework. The experimental programme became the source material for productised consulting work.
06 · Lessons

What I’ve learned about running personal labs

Three reflections from running this for 12 months:

  1. The discipline of pre-registration is more important than the experiments themselves. Half the value of this programme isn’t the findings — it’s the habit of writing down predictions before running tests. The number of times I’d have rationalised noise into “insight” if I hadn’t pre-committed is genuinely uncomfortable to admit.
  2. Personal labs compound differently than client work. Client engagements give you depth on a single system. A personal lab gives you breadth across patterns that surface only when you can run experiments without commercial pressure. The two are complementary; neither replaces the other.
  3. Most of what gets shared on AI search is wrong, but not maliciously. It’s wrong because the field is moving faster than rigorous experimentation can keep up. The honest thing is to keep running experiments, keep publishing what you find, and keep being explicit about confidence levels.

This work is ongoing. New cohorts ship each quarter. The findings will continue to evolve as the retrieval models do. The framework is more durable than the specific results.

Other case studies