← Back
technical seogoogle search consoleindexinginternal linkingcrawl budget

Discovered Currently Not Indexed: A 60,000-URL Recovery Timeline

A staged, evidence-led plan for moving a large site beyond “Discovered – currently not indexed” after repairing internal linking, without expecting an instant Search Console change.

· 15 min read

A site with roughly 700,000 content URLs had only 1,500 genuinely published and internally linked pages, while Google Search Console reported more than 60,000 URLs as “Discovered – currently not indexed.” After the site added a crawlable browse hub and footer links, the practical question was not whether Search Console would change tomorrow, but how to prove Google had received stronger discovery signals before the report caught up.

For the focus keyword discovered currently not indexed, the useful payoff is a recovery process that separates three events: Google recrawling changed pages, Google deciding whether a URL merits indexing, and Google Search Console reporting those changes. Those events do not occur on a single deadline. The scenario and numbers originate from a TechSEO community case, and they illustrate why a 60,000-URL issue on a 700,000-page system needs cohort measurement rather than a wait-and-refresh approach. (reddit.com)

Adding server-rendered internal links is a meaningful fix when important pages were effectively orphaned from Google's crawlable entry points. It does not, however, guarantee that the “Discovered – currently not indexed” total will start falling within a fixed number of days.

Google must first revisit the pages where the new links exist, discover the destinations through those links, decide how much crawl capacity to allocate, fetch destination URLs, assess canonicalization and content value, and then refresh Search Console reporting. Google explicitly says to allow at least a week after submitting a sitemap or indexing request before assuming there is a problem. For a large site, that is a minimum troubleshooting window, not a promise that 60,000 URLs will all be crawled or indexed in a week. (support.google.com)

A practical expectation is:

  1. Day 1: verify that Google can render and access the new discovery path.
  2. Week 1: look for proof of recrawls and changed discovery signals in a representative URL cohort.
  3. Weeks 2–4: assess direction of travel in the cohort and the Page Indexing report, then decide whether the remaining issue is crawl discovery, technical eligibility, or page quality.

The important standard is not “every discovered URL becomes indexed.” Google says important canonical pages should be indexed, while some non-indexed pages are expected and appropriate, including duplicates, removed URLs, or pages intentionally excluded from crawling or indexing. (support.google.com)

Why 700,000 URLs but 1,500 published pages changes the diagnosis

The headline number—60,000 discovered URLs—sounds like a simple internal-linking failure. The ratio behind it makes the diagnosis more complicated. If a platform can generate around 700,000 content pages but only about 1,500 meet a publication threshold and receive normal navigation links, Google may be seeing a very large universe of possible URLs without strong evidence that most are intended as durable, useful landing pages.

An XML sitemap tells Google that URLs exist, but it does not make every URL eligible for crawling or indexing. Google says a sitemap can be fetched immediately after submission, while crawling the listed URLs can take time and not all sitemap URLs will necessarily be crawled; site size, activity, traffic, and other factors affect that outcome. (support.google.com)

That distinction matters in this case:

  • The 1,500 published pages should be the first cohort to protect and measure. They have passed the site's own quality threshold and should have clear, crawlable internal links.
  • The other generated URLs need an explicit policy. Some may deserve publication and links; others may belong outside the sitemap, outside browse navigation, or behind an intentional noindex rule if they are thin, duplicate, transient, or not intended for search.
  • The 60,000 discovered URLs should not be treated as one homogeneous backlog. Some could be valuable pages that lack crawl paths. Others may be URLs Google found through historic sitemaps, parameter combinations, prior links, or generated structures that should never be a priority for indexing.

Before expanding internal links across hundreds of thousands of URLs, the site should define which URL classes are indexable assets. This is closely related to the question of what to fix first in a large audit: critical crawlability and indexability defects should precede cosmetic on-page work. See technical SEO audit prioritization for a practical way to rank such work.

What “Discovered – currently not indexed” actually says

“Discovered – currently not indexed” means Google knows about the URL but has not crawled it yet. It is therefore primarily a crawl and discovery status, not a final verdict that Google reviewed the page and rejected it.

That differs from “Crawled – currently not indexed.” In the second status, Google has fetched the page and has not added it to the index at that time. The likely investigation shifts: instead of asking only whether Google can find the URL, the audit must examine canonical selection, duplication, content distinctiveness, soft errors, and whether the page offers enough unique value to index.

The two statuses can occur on the same site and require different first moves:

Search Console statusWhat Google has doneFirst diagnostic priority
Discovered – currently not indexedFound the URL, but not crawled itCrawlable internal links, sitemap scope, server capacity, URL prioritization
Crawled – currently not indexedCrawled the page, but did not index itCanonical, duplication, content quality, page purpose, soft-404 signals

For any specific URL, Google recommends URL Inspection rather than relying on the aggregate Page Indexing report. The tool shows what Google knows about the indexed version, while a live test can check current crawl allowance, page fetch success, and whether indexing is allowed. (support.google.com)

The Reddit case identified an important technical pattern: content navigation existed after client-side hydration, but the homepage's served HTML contained no crawlable content links. The repair added a browse or index hub with real HTML links and connected it from every page footer. That is directionally sound because Google describes good navigation as a site where links can be followed from page to page until all pages can be found. (reddit.com)

The audit should verify the implementation rather than assume a front-end change solved it.

Check the root and hub pages

For the homepage, browse hub, category pages, and a sample of content pages, inspect the response that a crawler receives:

  • Confirm the links appear in the initial, served HTML—not only after a browser runs JavaScript.
  • Confirm destinations use normal crawlable <a href> links rather than click handlers, hash routes, or form submissions.
  • Confirm every important content cluster has a path from the homepage or a consistently linked hub.
  • Confirm links are not marked nofollow internally without a deliberate reason.
  • Confirm pagination, category filters, and faceted paths do not produce uncontrolled URL expansion.

A page linking only to peer pages can still be difficult for Google to prioritize if no durable, prominent route brings crawlers into that cluster. A hub linked from the footer may be useful, but an audit should also consider whether the hub is meaningfully reachable from primary navigation and whether its link architecture is shallow enough for key content.

For a broader internal-link review, Audra’s comparison of Semrush unoptimized anchors and internal-link review explains why raw link counts alone are not enough: destination coverage, link placement, anchor relevance, and orphan detection all matter.

Run the eligibility checks before waiting for a crawl budget response

A stronger crawl path cannot overcome a URL that is blocked, broken, non-canonical, or intentionally excluded. Before interpreting a slow result as a Google crawl-budget issue, test a sample from each URL class.

Use URL Inspection on roughly 20–50 URLs per cohort, rather than selecting random pages from one template only. Include recently linked pages, old discovered URLs, pages from different categories, and a small number of high-value pages.

For each inspected URL, record these fields:

  1. HTTP response: a stable 200 OK is expected for pages intended to index.
  2. Crawl allowed: this should be “Yes”; robots.txt blocking prevents Google from crawling the URL.
  3. Page fetch: this should be “Successful.”
  4. Indexing allowed: this should be “Yes”; a noindex directive prevents index inclusion.
  5. User-declared and Google-selected canonical: a mismatch may be legitimate, but it must be understood.
  6. Last crawl date: this establishes whether Google has revisited the page after the linking change.
  7. Referring sitemap: confirm that only intended canonical URLs are submitted.

Google’s documentation specifically identifies crawl permission, successful page fetch, indexing permission, and Google-selected canonical as core fields for diagnosing why an individual URL is absent from Google. (support.google.com)

If canonical URLs, robots directives, and HTTP responses vary by template, treat that as a template-level defect. Do not spend weeks waiting for Search Console to improve when the same technical obstacle applies to thousands of URLs.

Build cohorts instead of monitoring one 60,000-URL total

Search Console’s aggregate counts are useful for trend monitoring but weak as a short-term experiment dashboard. The report shows totals as of dates on its chart, and its example URL list is limited; it is not a complete real-time ledger of every affected URL. (support.google.com)

A cohort approach creates evidence. Create a spreadsheet or database with stable URL samples and a clear label for the discovery intervention applied.

Suggested cohort design

CohortSizePurpose
A: newly linked, published pages100 URLsTests the repaired browse hub and footer links on the 1,500 intended pages
B: high-value old discovered pages100 URLsTests whether previously known URLs receive new crawl attention
C: pages not linked or not published50 URLsControl group; helps distinguish broad Google changes from the new internal-link effect
D: likely low-value/generated pages50 URLsTests whether the quality threshold, not discovery, is the true constraint

Capture the baseline on the day of deployment: URL Inspection status, last crawl date, canonical, sitemap inclusion, crawlability, and whether the source HTML contains the intended link path. Recheck the same URLs at day 7, then at days 14 and 28.

The signal to seek is not instant indexing. It is a difference between cohorts. If Cohort A shows new last-crawl dates and more “URL is on Google” outcomes than the control group, the new internal links are likely being noticed. If Google recrawls Cohort A but leaves many URLs crawled-not-indexed, the work should shift toward content and canonical diagnosis.

A 1-day, 1-week, and 2–4-week measurement plan

The recovery plan should use fixed checkpoints so teams do not confuse normal reporting delay with a failed fix.

Within 24 hours

  • Fetch the homepage and browse hub as raw HTML and confirm real destination links are present.
  • Inspect 10 URLs from Cohort A with URL Inspection’s live test.
  • Confirm the XML sitemap is accessible and contains only canonical URLs intended for indexing.
  • Submit or resubmit the sitemap if it changed. A successful sitemap fetch is useful confirmation, but it is not confirmation that all URLs were crawled. (support.google.com)
  • Request indexing for the homepage, browse hub, and a small representative sample of high-priority content pages.

At day 7

  • Reinspect the same sample and compare last crawl dates with the baseline.
  • Review server logs, if available, for Googlebot hits to the hub and linked destinations.
  • Compare crawl activity for Cohort A with the control group.
  • Review the Page Indexing report trend, but do not require the 60,000 figure to have materially changed yet.

At weeks 2–4

  • Measure how many URLs in Cohort A were recrawled, indexed, still discovered-not-indexed, or moved to crawled-not-indexed.
  • Segment results by template, category, publication date, depth, and content type.
  • Remove non-indexable generated URL classes from sitemaps where appropriate, rather than asking Google to process them repeatedly.
  • Escalate technical defects only where cohort evidence identifies a consistent pattern.

This time-boxed process prevents an unhelpful conclusion such as “Google ignores the site.” It replaces it with a testable statement—for example, “Google discovered the new hub within seven days, crawled 62 of 100 newly linked published URLs, and withheld indexing mainly on a duplicated template.” The exact numbers will vary, but the measurement logic does not.

Do not request indexing for all 60,000 URLs

Google positions the URL Inspection request feature as a way to request crawling for a URL, and its help documentation discusses it in the context of individual-page troubleshooting. It should be used for representative or business-critical URLs, not as a substitute for scalable architecture. (support.google.com)

For this scenario, request indexing for:

  • the homepage;
  • the new browse or category hub;
  • a handful of important category pages;
  • 10–20 representative published content pages; and
  • a small number of previously discovered high-value URLs.

Avoid trying to force all 60,000 through individual indexing requests. That produces little diagnostic insight and can hide the real question: whether Google can consistently discover, crawl, and choose to index the site’s intended canonical pages at scale.

If the site's quality threshold means only 1,500 pages are intentionally published, the sitemap should reflect that editorial decision. Sending hundreds of thousands of marginal URLs through a sitemap can make prioritization less clear, not more effective.

When internal linking is fixed but pages still do not index

If the new links are visible in served HTML and Cohort A gets recrawled, the original orphaning problem may be resolved even if the aggregate report remains high. The next diagnosis depends on the new URL status.

If pages remain “Discovered – currently not indexed”

Review site scale, URL count growth, sitemap scope, duplicate parameter URLs, and whether Googlebot can efficiently traverse the hub. A site with 700,000 possible content pages needs a tighter definition of what should be crawlable and indexable than a 500-page brochure site. Google itself recommends the Page Indexing report for larger sites, while stating that good navigation should let crawlers find all pages through links. (support.google.com)

If pages become “Crawled – currently not indexed”

Investigate content uniqueness and canonical intent. Compare pages within a template for duplicated titles, near-identical copy, thin records, empty states, weak differentiation, and canonical tags pointing elsewhere. The fix is often not more links; it is deciding which pages offer standalone search value.

If important pages index but rankings remain weak

Indexing and ranking are separate outcomes. Once indexing is stable, review search intent, page differentiation, internal anchor context, and whether multiple pages compete for the same query. For B2B sites, home page keyword cannibalization shows why a homepage can unintentionally compete with category pages even after crawlability is fixed.

Use an audit to validate the whole system, not just one report label

A “Discovered – currently not indexed” recovery can fail quietly when teams inspect only Search Console totals. The better audit combines crawlability, indexability, internal links, canonical behavior, performance, and the HTML that bots actually receive.

For example, a local crawl can identify orphaned content pages, links missing from rendered HTML, redirect chains, noindex directives, non-canonical sitemap entries, and pages that require JavaScript to expose core navigation. Audra can be used to crawl those relationships locally alongside technical SEO, performance, accessibility, and link checks, which makes it easier to compare the intended published set with the pages actually reachable through HTML links.

The goal is not to manufacture a perfect indexation rate. It is to make sure the business-critical, canonical, quality-approved pages are easy for Google to find and evaluate, while low-value generated URLs do not consume disproportionate attention. For a related diagnostic framework, see Discovered – Currently Not Indexed: a guide to finding the real blocker.

FAQ

How long does Google usually take to recrawl pages after internal linking is fixed?

There is no reliable fixed timeline. Google advises allowing at least one week after a sitemap submission or indexing request before assuming a problem, but large sites can require longer and not every URL will be crawled. Start checking a stable cohort after seven days, then compare outcomes at 14 and 28 days rather than waiting for one report total to change. (support.google.com)

No. New internal links can improve discovery and crawl prioritization, especially for pages that were orphaned from crawlable HTML. Google still needs to crawl each URL and decide whether it is canonical and worth indexing. If a page later becomes “Crawled – currently not indexed,” investigate duplication, canonicalization, and content quality rather than adding more links.

What should a site with 60,000 discovered URLs and 700,000 total content pages do first?

First, define the set of pages that are genuinely intended to index. In the example, only about 1,500 pages had passed a publication threshold and were internally linked, so those pages should become the priority cohort. Verify crawlable navigation, canonical status, robots directives, HTTP responses, and sitemap inclusion before attempting to push all 60,000 URLs into Google's crawl queue.

Use the same 100 or so newly linked pages as a tracked cohort. In URL Inspection, compare last crawl dates before and after deployment, and check whether Googlebot visits the hub and destination pages in server logs when logs are available. A higher recrawl rate for newly linked pages than for a control group is stronger evidence than a short-term change in the aggregate Search Console count.

Should a site request indexing for all affected URLs or only a representative sample?

Use requests for the homepage, new hubs, critical category pages, and a representative sample of important content URLs. Google documents the feature as an individual-URL troubleshooting tool, while sitemaps and internal links are the scalable discovery mechanisms. Mass requests for 60,000 URLs do not solve weak site architecture or quality problems. (support.google.com)

Sources