How to Read a Site's History on the Wayback Machine Before You Buy It
The archive is the one record a seller cannot edit. How to sample snapshots properly, detect niche switches, diff historical robots.txt files, and query the CDX API to see what a site really was.
Checklist item 002: Check the site's history on Wayback Machine: no radical niche changes
Open checklist →Every other number in a due diligence pack passes through the seller's hands first. Analytics can be filtered, revenue dashboards can be screenshotted at a flattering moment, and a backlink profile can be described as "clean" by someone who has never audited it.
The Internet Archive is different. It was recording the site years before anyone decided to sell it, it was not curated for your benefit, and the seller cannot go back and change it. Used properly it will tell you what the site actually was — which is frequently not what the listing says.
Why the Archive Is the One Source the Seller Can't Edit
A site's history determines what you are really buying. Two sites with identical traffic and revenue today can be worth very different amounts:
- One has published in the same niche for seven years, accumulating topical authority that Google has had seven years to trust.
- The other was a coupon site, then a crypto blog, then a dropshipping store, and became a "health and wellness authority site" fourteen months ago.
The second one is a fourteen-month-old site wearing a seven-year-old domain. You should value it as the former and pay for it as the former, no matter what the WHOIS record says.
Niche switches matter because Google's trust is topical, not just domain-wide. When a site changes subject entirely, the authority it built in the old niche largely does not transfer — and the backlinks pointing at the old content become, at best, irrelevant, and at worst a signal that something odd happened here.
This check also feeds directly into two others. If you find a gap or a switch, you will want to know whether it coincided with a Google penalty or an algorithmic hit, and whether the old content was the kind of thing that contaminates a domain permanently.
Step 1: Read the Calendar Properly
Go to web.archive.org and enter the bare domain.
The default view is a bar chart of captures per year, and below it a calendar for the selected year. Most buyers glance at this and conclude "lots of blue dots, looks fine". That is not reading it.
What the density actually tells you. Crawl frequency is roughly proportional to how much the wider web cared about a site. A domain that was genuinely popular in 2019 gets crawled often in 2019. So the shape of the bar chart is a crude, independent traffic history — one the seller had no hand in.
Compare that shape against the traffic story you have been told. A seller claiming the site has grown steadily for five years, on a domain whose capture density peaked in 2018 and thinned out afterwards, is describing a different site from the one in the archive.
What to look for:
- Density that rises and then collapses — the site lost relevance at a specific point. Find that point and ask what happened.
- Multi-year gaps — the site was down, parked, or blocking crawlers. All three matter.
- Density that only begins recently on a supposedly old domain — the domain existed, but the site did not. Cross-check against the domain's registration history.
- Suspiciously uniform density — occasionally a sign the owner was submitting pages to the archive themselves rather than the crawler finding them naturally.
Step 2: Sample Snapshots on a Schedule, Not at Random
Clicking a few dots at random is how buyers miss things. Sample deliberately:
- The earliest available capture. What was this domain originally for?
- One capture per year, ideally the same month each year, for the whole history.
- Two captures either side of every gap you found in Step 1 — the last one before, the first one after. This pair is where niche switches and ownership changes reveal themselves.
- One capture immediately before any traffic drop the seller has disclosed.
For each one, record four things: the niche, the site name and branding, the monetisation visible on the page, and roughly how much content is on the homepage.
That last one matters more than it sounds. A homepage that showed forty article links in 2021 and shows nine in 2025 tells you content was removed — and removed content is often removed because it was penalised, thin, or plagiarised.
Step 3: Identify Niche Switches and Rebrands
A clean site tells one continuous story. A repurposed domain tells several.
The tells, in rough order of severity:
| What you see across snapshots | What it usually means |
|---|---|
| Same brand, same niche, content growing | A genuine long-running site |
| Same niche, new design and brand name | A normal rebrand — verify the content survived the redesign |
| Related niche shift (personal finance → investing) | Usually fine; topical authority partly transfers |
| Unrelated niche shift (travel → CBD) | The domain's age is cosmetic; value it from the switch date |
| Multiple unrelated niches over time | A recycled domain that has been flipped before |
| Foreign-language content that later became English | Old backlinks are from the wrong language and market |
| Parking pages ("this domain is for sale") for long stretches | No organic history was built during those years |
A specific pattern worth knowing: an expired domain with strong backlinks gets bought at auction, rebuilt in a completely different niche, aged for twelve to eighteen months, and sold as an "established authority site". The archive is the only place this shows up plainly. The WHOIS record may have been privacy-protected throughout, the analytics only go back as far as the rebuild, and the backlink profile looks great precisely because it belongs to a site that no longer exists.
Step 4: Diff the Historical robots.txt
This is the step almost nobody performs, and it is often the most revealing.
robots.txt is archived like any other file. Open the history for it directly:
https://web.archive.org/web/*/example.com/robots.txt
Then compare versions across the years. You are looking for Disallow rules that existed once and are gone now — because they name the parts of the site the previous owner did not want crawled.
What findings mean:
Disallow: /across a long period — the site was deliberately hidden from crawlers, including the archive. That explains a gap, but it does not excuse it. Ask why.- A blocked directory that no longer exists — go and see what was in it.
/go/,/out/,/recommends/and similar are typically affiliate cloaking or link redirection. - A rule blocking
/tag/,/page/or search pages — entirely normal SEO hygiene, not a finding. User-agent: ia_archiverwithDisallow: /— an explicit instruction to keep the Internet Archive out. On a site being sold as transparent, that is worth a direct question.
Apply the same technique to sitemap.xml. URLs that appear in old sitemaps and are absent from the current one are pages that were deleted. A few is editorial pruning. Several hundred is a cleanup, and you want to know what was being cleaned.
Step 5: Query the CDX API for the Whole Picture
Clicking through snapshots does not scale past a few dozen. The Internet Archive exposes a query interface — the CDX Server API, one of the official Wayback Machine APIs — that returns the capture index as plain text or JSON, and it turns this check from an afternoon into a few minutes.
The base endpoint is http://web.archive.org/cdx/search/cdx, and url is the only required parameter. The publicly available fields are urlkey, timestamp, original, mimetype, statuscode, digest and length.
Three queries do most of the work.
One row per year, to see the shape of the history:
https://web.archive.org/cdx/search/cdx?url=example.com&fl=timestamp,statuscode,mimetype&collapse=timestamp:4&output=json
collapse=timestamp:4 keeps only the first capture per distinct four-digit prefix — that is, per year. Use timestamp:6 for monthly.
Everything that wasn't a normal page, to find parking and redirects:
https://web.archive.org/cdx/search/cdx?url=example.com&filter=!statuscode:200&output=json
Runs of 301 and 302 mean the domain was redirecting somewhere else — find out where, because that period built no authority for this site. Long stretches of 404 mean it was dead.
Every URL ever captured on the domain, to see its true size and shape:
https://web.archive.org/cdx/search/cdx?url=example.com&matchType=domain&fl=original&collapse=urlkey&output=json
This is the one that catches things. Subdomains you were never told about, an old /shop/ section, a forum that was quietly removed, thousands of near-identical URLs that suggest programmatic content. Compare the count against the number of pages the seller says the site has.
The digest field is a content hash, so identical digests across consecutive captures mean the page did not change at all between them. A homepage with the same digest for two years was not being maintained, whatever the listing says about "regularly updated content".
A practical note: the archive rate-limits aggressively. Request politely, one query at a time, and expect 429 responses if you loop over URLs quickly.
Step 6: Look for the Ownership Transition
Somewhere in the history there is usually a moment where the site changed hands. Find it, because everything before it belongs to someone else's decisions and everything after it is what you are actually buying.
The markers cluster together: the design changes wholesale, the author bylines change or vanish, the "About" page is rewritten in a different voice, the monetisation switches (display ads appear, or an affiliate network changes), and the publishing cadence shifts.
Compare that date against what the seller told you. "I've been running this site since 2019" is a different claim from a site whose entire character changed in 2023. Sellers are not always lying about this — plenty bought the site themselves and count from the domain's birth out of carelessness — but you need the real number, because the post-transition period is the only part whose performance predicts anything.
What to Do With What You Find
| Finding | Action |
|---|---|
| One continuous niche, capture density consistent with the traffic story | ✅ Proceed |
| Clean rebrand within the same niche, content preserved | ✅ Proceed |
| Related niche shift more than two years ago, stable since | ⚠️ Verify rankings held through the shift |
| Gap of 1–2 years with a plausible, verifiable explanation | ⚠️ Treat the post-gap period as the site's real history |
| Unrelated niche switch | ⚠️ Value the site from the switch date, not the registration date |
Old Disallow rules hiding directories that no longer exist | ⚠️ Investigate what was there before going further |
| Hundreds of URLs in old sitemaps, absent now | ⚠️ Ask what was removed and why |
| Long periods of parking pages or redirects to another domain | ⛔ The domain's age built no authority for this site |
| Several unrelated niches over the domain's life | ⛔ A recycled domain that has been flipped before |
Archive blocked entirely (ia_archiver disallowed) for long stretches | ⛔ You cannot verify the history, so you cannot price the risk |
| History flatly contradicts what the seller told you | ⛔ Walk away — this is a disclosure problem, not a data problem |
Quick Reference Checklist
Before signing off on this item:
- Capture-density chart reviewed across the domain's full life and compared against the seller's traffic story
- Earliest capture examined — you know what the domain was originally for
- One snapshot per year sampled, plus pairs either side of every gap
- Niche, branding, monetisation and content volume recorded for each sample
- Any niche switch dated, and the site revalued from that date
- Historical
robots.txtversions diffed for removedDisallowrules - Old
sitemap.xmlversions compared against the current one for deleted pages - CDX query run for non-200 status codes — no unexplained redirect or parking periods
- Full URL list pulled and reconciled against the seller's claimed page count
- Ownership transition dated and checked against the seller's account of it
This takes thirty to forty minutes and needs no paid tools. It belongs early in the due diligence checklist, right alongside the domain age check — because if the history and the story do not match, nothing you verify afterwards is worth the time.