An open dataset on what AI engines actually cite — including the parts that don’t flatter me.
Two websites. One owner. The same 3 months, measured the same way. 21,700+AI citations between them — and completely opposite results. Everything below is free to download, re-use and check, under CC BY 4.0.
Same owner, same window, opposite outcomes.
This is the finding, and it is the reason the dataset is worth publishing rather than just the total. Both sites were read from the same console on the same day, on the same trailing window.
Canadian real-estate calculators. Purpose-built, brand new. Under three months old at the start of the window. Curve: near-zero → steep, sustained growth.
A personal blog that later became a consultancy site. About 18 months old at the start of the window. Curve: flat across the entire window — no growth, ~91% concentrated on a single page.
Roughly comparable totals. One is spread across 25+ pages and climbing; the other is dominated by a single page and flat. A citation total, on its own, tells you almost nothing.
| File | What it holds |
|---|---|
| 01_site_summary.csv | Site-level totals, pages earning citations, growth-curve shape, for both sites. |
| 02_homecalc_most_cited_pages.csv | The six most-cited pages on HomeCalc.ca, typed as Tool or Guide. |
| 03_homecalc_top_queries.csv | Top grounding queries with citation counts and citation share. |
| 04_hamitahm_most_cited_pages.csv | The six most-cited pages on HamiTahm.com — including the ones that embarrass me. |
| 05_commercial_reality.csv | Citations vs. actual Google clicks for the single most-cited page. |
| METHODOLOGY.md | Source console, window, pull date, what a citation is, how rounding is handled. |
| LIMITATIONS.md | Six stated limits, including n=2 and single-engine coverage. |
| DATA_DICTIONARY.md | Column-by-column definitions for every CSV. |
The selected-sample CSV is served directly from this domain and will always resolve: download it here. For the complete per-page export, or the underlying console screenshots, email hami@hamitahm.com — I send them.
Tahm, H. (2026). AI Citation Study: Two Sites, One Owner, Three Months [Data set]. Zenodo. https://doi.org/10.5281/zenodo.21651568
Archived copy: Zenodo record.
Every figure was read off a console screen.
Source: Bing Webmaster Tools → AI Performance (Microsoft Copilot and partners). Window: April 25, 2026 – July 25, 2026. Pulled: July 27, 2026. Click and position data comes from Google Search Console over the same window. Nothing here is modelled, projected, or rounded upward.
Where the console displayed its own rounded figure — “1.7K” rather than an exact integer — the dataset publishes that same rounded string, so it never implies more precision than the source gave. The full measurement method, including how prompts and conditions are recorded for the non-Copilot engines, is on the methodology page.
Stated up front, not buried in a footnote.
A strong signal, not a law. Two data points cannot support general claims about how AI citation works — only about what happened to these two sites, in Canada, in this window.
This measures Microsoft Copilot and its partners, because that is the only engine that reports citation counts back to publishers. Google Search Console has since added a Generative AI features report, but it gives impressions only — not citations — so it cannot be compared to, or added to, these figures. ChatGPT, Gemini and Perplexity expose nothing. Behaviour there may differ.
Both sites share an owner, so there were no competing stakeholders and no legacy debt on the newer one. A business without that control should expect a slower, messier curve.
The dataset shows what got cited, how much, and how fast. It does not explain how the pages were built to earn it — that is the paid work, and saying so plainly is more honest than pretending the dataset is a full recipe.
The most-cited page in the entire study (6,500 AI citations) produced 24 Google clicks over the same three-month window. That file is included on purpose.
The tables show the leaders, not every page or query that received a citation. The complete per-page export is available on request.
Grounding queries are Bing's own aggregated groupings of prompt activity, not verbatim user prompts, and Bing itself describes AI Performance as sampled, aggregated reporting rather than a complete log. This dataset publishes those figures faithfully — it does not and cannot de-aggregate them.
Use this data. Please.If you are writing about AI search and want real numbers instead of speculation, take them — charts, tables, figures. The licence is CC BY 4.0: commercial use included, attribution the only condition.
Credit “Hami Tahm” and link back to this page so readers can check the source themselves.
Want this run on your own site?
The same measurement, on your queries and your competitors — dated, attributed, and honest about what it can’t tell you.