A page can be published, work perfectly in your browser and still be difficult for a search engine to find. Publishing creates a URL; discovery gives a crawler a way to learn that the URL exists. Keeping those two events separate makes technical problems much easier to diagnose.
Consider a newly published guide in a resource library. If it appears in a relevant category and another guide links to it, both readers and crawlers have routes to the page. If its only entry point is a search form, the URL may be much harder to discover.
Discovery, crawling and indexing are different steps
Discovery means learning about a URL. Crawling means requesting it and retrieving its response. Indexing means processing the retrieved material and deciding how it belongs in a searchable collection. A successful crawl does not guarantee indexing, and an indexed page is not guaranteed to appear for a particular query. Google’s explanation of how Search works describes these stages separately.
That distinction changes the next action. An undiscovered URL calls for better discovery paths. A known URL returning an error calls for a server or routing fix. A fetched page excluded as a duplicate calls for an examination of content and canonical signals.
Links create discoverable routes
A normal link has a destination in an anchor element’s href attribute. A visitor can follow it, and a crawler can extract the destination. Interactive controls that only navigate through a script need closer inspection. Do not assume that a button which works when clicked is an equivalent discovery mechanism.
Build routes around a reader’s next question. A category can list its guides; a guide can point to a prerequisite or a more detailed explanation. Our site architecture guide explains how those routes form a maintainable hierarchy. A site with coherent routes is easier to audit than a collection of pages linked only from a sitemap.
Sitemaps help communicate your URL inventory
An XML sitemap supplies a list of URLs you want a search engine to know about. Treat it as an inventory of preferred, indexable pages. Include the final URLs rather than redirecting addresses, private pages or deliberately excluded duplicates. When your publishing system generates sitemaps, check that newly published content appears and removed content disappears.
A sitemap is useful evidence during an investigation, but submission is not a promise of crawling or indexing. Compare three inventories: published content, URLs reached through internal links and URLs in the sitemap. Differences identify pages worth inspecting.
Robots.txt controls crawling, not confidentiality
A robots.txt rule asks compliant crawlers not to request certain paths. It is not access control, and it is not a dependable way to remove a URL from search results. A URL blocked from crawling can still be known through links. A crawler also needs access to a page to read a noindex directive placed in that page. See Google’s robots.txt guidance.
For confidential material, require authentication. For a public page that should be excluded from search, choose an appropriate indexing control and verify that a crawler can read it. Avoid adding broad robots rules merely because a crawler tool reports many URLs.
Find orphan pages before adding more content
An orphan page has no discoverable internal links from the site’s accessible page network. It might still receive visits through bookmarks or external links, so a traffic report alone will not expose the problem. Export published URLs from the CMS and compare them with a crawl started from the homepage.
For example, if a guide is published and listed in the sitemap but absent from the crawl, inspect its category assignment, pagination and contextual links. Add it to the right hub and connect it from a useful related page. The internal linking guide gives a repeatable way to choose those links.
A practical crawlability check
- Open the final URL and confirm a successful response with the expected content.
- Trace at least one ordinary link path from a relevant hub.
- Check whether robots rules allow access to the page and essential resources.
- Compare the sitemap entry with the preferred URL.
- Inspect indexing directives and canonical signals separately from discovery.
- Use available search-engine inspection reports or verified crawler logs to test the diagnosis.
Working principle: diagnose the stage that failed before changing the site. A useful page needs a reliable route, a retrievable response and consistent indexing signals. Use the technical audit checklist to extend this check across page templates.