Google does not see your site
What we are solving
Section titled “What we are solving”The site has been live for weeks and search returns nothing for it. Not a weak position — no result at all.
Indexing fails from the top down. First a crawler has to fetch the page, then crawl it, then index it, and only after all of that can the page rank.
A rewritten title on a noindex page changes nothing. So the checks below run in order: stop at the first one that fails, because while the page is blocked you cannot measure anything else on it.
Does the crawler get any text at all
Section titled “Does the crawler get any text at all”The first thing I check is what the server hands over before any JavaScript runs. Fetch the page and look for a sentence you can read on it with your own eyes:
curl -sL https://example.com/page | grep -i "a sentence you can see on the page"If grep says nothing and the browser shows the text, the browser is drawing it: such an app answers with a near-empty <body>, and the crawler gets a page with nothing to index. You fix that with server rendering or a prerender step at build time. A tag will not fix it.
What is actually in your robots.txt
Section titled “What is actually in your robots.txt”Pull the file from the live domain and read all of it — not from memory, and not only the lines you wrote yourself:
curl -s https://example.com/robots.txtMost often it is a Disallow: /: on a draft domain that line belonged there, and it rode into production with the rest of the config. The other failure is quieter: Google reads only the first 500 KB of the file, a generated robots.txt can truncate silently in the middle, and everything below the cut does not exist for Google.
Are you asking search to drop the page yourself
Section titled “Are you asking search to drop the page yourself”noindex lives in two places. The meta tag sits in the HTML, where you can see it in View Source, and the X-Robots-Tag response header never shows up there at all:
curl -sI https://example.com/page | grep -i x-robots-tagcurl -sI sends a HEAD request and prints the headers alone, and in a browser the same thing is under Network, the document request, Response Headers.
That header can come from the framework, the web server or the CDN. Those are three different files, and you will have to open each one. There is also a per-agent form — X-Robots-Tag: googlebot: noindex — and it looks like an ordinary line of config, so the eye slides right past it.
Search fetches a noindex page and reads it honestly enough. Then it drops the page on purpose.
How many addresses the page has, and where it points
Section titled “How many addresses the page has, and where it points”Every page should carry a canonical pointing at itself or at the real original:
curl -sL https://example.com/page | grep -i 'rel="canonical"'A bad template sets the canonical to the home page on every page at once. That is a request to search: drop everything except the home page.
Then count how many addresses lead to the same page. Trailing slash, www, http, index.html, tracking tags — each one is another address. All of them have to redirect to one:
for u in http://example.com/page https://www.example.com/page https://example.com/page/ https://example.com/page/index.html; do curl -o /dev/null -sw "%{http_code} %{redirect_url}\n" "$u"doneOtherwise the signal splits between duplicates and the sitemap starts disagreeing with the canonical tag. Search settles that argument on its own, without you.
Is the sitemap alive, or merely present
Section titled “Is the sitemap alive, or merely present”Do not check that the file exists. Check that every URL in it answers 200 and that each one is written out in full, in the canonical form:
curl -s https://example.com/sitemap.xml | grep -o '<loc>[^<]*' | cut -c6- | while read -r url; do curl -o /dev/null -sw "%{http_code} $url\n" "$url"; doneThe protocol caps one file at 50,000 URLs and 50MB uncompressed. Past that, split the file and add a sitemap index, and if a map is stuffed with redirects and 404s, it stops being read as the truth about the site.
Is the site visible from outside your network
Section titled “Is the site visible from outside your network”Do not run the last check from your own machine. Resolve the domain from someone else’s address and request the page without cookies:
ssh other-box 'curl -sI https://example.com/page'Access control, basic auth and edge bot rules answer with a login page or a 403. None of that is visible in robots.txt.
The edge rule has a name and a screen it lives on: on Cloudflare it is Configure AI bot policies, on the zone’s Security Settings page, on every plan. Cloudflare refuses a matching agent with a 403 from its own network, and that request never reaches your server.
So the refusal is not in your server log either: you can see it only in Cloudflare’s Analytics → Events. The whole mechanism is on AI crawlers and llms.txt.
What did not work
Section titled “What did not work”- Waiting for the index to catch up. The crawler does come back, reads the same rule and leaves again. Patience does not edit a header.
- Re-submitting one URL in the inspection tool. The re-fetch uses the same
robots.txt, the same header, the same empty body. The verdict comes back identical. The daily quota is gone. - Trusting a third-party crawler as proof of access. The report says what someone else’s bot could fetch from its own address, and whether search decided to keep the page is a different question.
- Checking only from my own laptop. A logged-in session, a warm service worker and a home network the edge already trusts hid the failure between them. The edge is the CDN in front of your origin — Cloudflare, Fastly, a load balancer — and it answers some requests before they reach you.
- Editing
robots.txtwhen the block lived at the edge. Bot-protection and WAF rules are invisible in that file, and the file was clean the whole time. Look for Cloudflare’s Configure AI bot policies under Security Settings, and any WAF custom rule beside it — this is the most common hidden blocker I run into. - Rewriting titles and descriptions first. On a page that is not in the index, on-page work produces nothing you can measure.
Verify
Section titled “Verify”Run the seo-audit skill from Tools: it walks this same order and collects the answers into one report, and doing the same by hand takes longer.
Then check by hand what the report cannot fake:
curl -sIon the page returns 200 and carries noX-Robots-Tag.curl -sLon the same URL contains a sentence you can read on the page.- URL Inspection in Search Console says the URL is on Google.
- Cloudflare’s Analytics → Events, filtered to the Block action, holds nothing for your own URLs, because a blocked crawler shows there and nowhere on your server.
To see roughly how many pages reached the index at all, ask the search box:
site:example.comRead the wording in Search Console literally. “Discovered — currently not indexed” means the URL is known and was not fetched, while “Crawled — currently not indexed” means it was fetched and judged not worth keeping. Those are two different bugs, and you fix them differently.
Once one page is in, hand over the rest: submit and verify.
