Complete Guide

How Do You Turn a Screaming Frog Crawl Into Better On-Page SEO Without Losing 6 Hours in Low-Value Reports?

Configure the crawl around your site, combine crawl evidence with search data, work through issues in dependency order, and re-crawl to confirm that each change solved the intended problem.

13 min read

Quick Answer

What to know about How to Use Screaming Frog for a Reliable On-Page SEO Audit

A useful Screaming Frog audit starts with a documented question, scope, rendering mode, and saved baseline. Review page purpose and duplication before architecture, navigation, orphan candidates, performance evidence, and final on-page elements.

Use custom XPath or CSS extraction to test template requirements, but validate empty or inconsistent results in rendered pages before assigning work. Combine crawl depth, All Inlinks, Google Search Console, and analytics data to form hypotheses about important URLs, then inspect canonical, indexability, content, and demand before choosing a fix.

Titles and meta descriptions should be revised from real query evidence and page purpose, not warning counts alone. Every recommendation should include an acceptance test, and every implementation should be confirmed with a targeted re-crawl followed by a comparable later review.

A Screaming Frog crawl is useful only when it helps you decide what to change, what not to change, and how to verify the result. The tool can expose response codes, canonicals, page titles, headings, content similarity, internal links, crawl depth, rendered HTML, and data imported from connected services.

It cannot decide whether a URL deserves to exist, whether two pages serve the same intent, or whether a flagged field is the reason performance is weak. That judgement belongs in the workflow.

This guide takes you from preparation to validation. You will define the crawl purpose, capture a baseline, configure rendering and scope, collect internal and external URL evidence, classify issues by dependency, implement changes in controlled batches, and compare a fresh crawl against the baseline.

The intended outcome is a prioritised on-page SEO worklist in which every recommendation names the affected URL, the evidence, the proposed action, the acceptance test, and the next step when the evidence is inconclusive.

Use the crawl as one input, not as an automatic instruction set. A duplicate title may reflect an intentional pagination pattern, a missing main heading may be supplied after rendering, a low word count may be appropriate for a utility page, and a deep URL may still be well supported through relevant links.

Conversely, a technically valid successful-response page can still be a poor search landing page because its purpose is unclear, its internal links are weak, or its canonical points elsewhere. The method below is designed to surface those distinctions before you spend time rewriting the wrong element.

Key Takeaways

  • 1Crawl before editing: establish a reproducible baseline, record the configuration, and avoid changing pages before you know whether the problem is content, architecture, or indexability
  • 2Use a 5-layer review sequence to separate content, architecture, navigation, orphan discovery, performance, and final on-page checks without treating every flag as equally urgent
  • 3Do not prioritise title tags automatically: first confirm that important URLs are crawlable, indexable, internally supported, and mapped to the right intent
  • 4Use custom extraction with XPath or CSS selectors to inspect headings, repeated template elements, JSON-LD presence, and page-specific fields at scale
  • 5Use crawl analysis and link reports to identify important pages that sit too deep, receive weak internal support, or are only discoverable through a sitemap or external data source
  • 6Treat low word counts as a review trigger, not a verdict: compare the page purpose, visible main content, duplication pattern, and search demand before deciding whether to expand or consolidate
  • 7Reconcile crawled URLs, sitemap URLs, and analytics or Search Console landing pages to find orphan candidates, then verify each candidate before adding links or removing it
  • 8Pair response codes with URL importance so broken internal targets, redirect chains, and unstable pages are fixed according to user and search impact rather than report order
  • 9Connect Google Search Console when available so impressions, clicks, and query evidence can help distinguish a discoverability problem from a snippet or content problem
  • 10Re-crawl with the same saved settings after implementation and compare results; a crawl is evidence for a decision, not proof that rankings or traffic will change

1Set the Audit Question, Scope, and Crawl Configuration Before You Start

The first decision is not which tab to open. It is what the crawl needs to prove or disprove. Write the audit question, the site area in scope, the date, the expected host and protocol, and the data sources you will compare.

Typical inputs are the start URL, XML sitemaps, a list of priority landing pages, Google Search Console, analytics data, and any known staging, parameter, or authentication rules. Keep a copy of the configuration and the original crawl file so the work can be repeated.

Start with Spider mode and confirm the correct hostname. Review Configuration > Spider for crawlable resource types, external link handling, robots directives, canonicals, pagination, and URL parameters.

Do not select Googlebot merely to make the crawl look more authoritative. Use the user agent that matches the question you are testing, and document it. If the site varies content by user agent, run a controlled comparison rather than assuming one crawl represents every search system.

For a site that relies on client-side JavaScript, compare 'Text Only' and 'JavaScript' rendering on a representative sample before committing to a full rendered crawl. Go to Configuration > Spider > Rendering, select 'JavaScript', and inspect the rendered HTML and screenshot for several templates.

Confirm that navigation, main content, canonicals, headings, and internal links appear after rendering. Rendering can reveal content that raw HTML does not contain, but it also increases crawl time and resource use.

Decide how canonicals should be handled. One pass can follow normal canonical behaviour to model the preferred crawl set; another can retain non-canonical URLs so you can inspect inconsistent signals, duplicate templates, and parameter variants.

Label the files clearly so findings are not mixed. If the result depends on canonicals, compare the declared canonical, status code, indexability, internal links, and sitemap inclusion before recommending a change.

Connect Google Search Console and Google Analytics through Configuration > API Access only when access is authorised and the chosen date ranges are understood. Imported metrics help prioritise URLs, but they do not prove causation.

A page with impressions and few clicks may need a title review, a better intent match, or no change at all if the queries are irrelevant. Record the selected property, dimensions, and date window beside the crawl.

Set limits that match the site. For most sites under 10,000 pages, a complete crawl may be practical. For a larger or unstable site, begin with a representative subfolder or template group, then expand.

If you use a crawl depth limit of five to seven levels, state that deeper URLs may be missing by design. Validate the setup by checking a small sample of known URLs before allowing the full crawl to run.

Write the audit question and expected decision before crawling so each report has a defined use
Choose the user agent deliberately and document it; compare user agents only when the site may serve materially different content
Test JavaScript rendering on representative templates before a full crawl and verify the rendered HTML, links, headings, canonicals, and main content
Use separate, clearly labelled crawl passes when you need to compare canonical-respecting and canonical-discovery behaviour
For large sites, segment by subfolder or template when a full crawl would be slow, unstable, or difficult to interpret
Save the configuration and baseline crawl so the validation crawl can use the same settings

2Review Findings in Dependency Order Instead of Fixing the Easiest Flags First

A crawl produces many observations at once, so use an ordered review that prevents downstream edits from being made before upstream constraints are resolved. The existing six-part sequence can be treated as a practical queue rather than as a claim that one factor always controls rankings. The purpose is to ask the right question at each stage and to stop when the evidence points elsewhere.

C - Content: Confirm that each important URL has a clear purpose, sufficient visible main content for that purpose, a coherent heading structure, and no unintended duplication. Use the Content tab, Exact Duplicates, Near Duplicates, and custom extraction.

A page with fewer than 300 words is not automatically thin; it is a review candidate. Compare it with the task the page must complete, the query set, the template, and nearby pages before choosing to expand, merge, redirect, noindex, or leave it unchanged.

A - Architecture: Run crawl analysis after the crawl is complete and inspect crawl depth, directories, click paths, and internal link distribution. A commercially important page at crawl depth five or deeper deserves review, but depth alone does not establish a problem.

Check whether the URL is linked from relevant hubs, whether users can reach it naturally, and whether the path reflects a deliberate information architecture.

N - Navigation: Export All Inlinks and separate sitewide links from contextual body links. Confirm that menus, breadcrumbs, related-content modules, and template links support the intended hierarchy.

If navigation heavily promotes low-value or obsolete URLs while priority pages rely on isolated body links, create a template or information-architecture task rather than editing page copy.

O - Orphans: Compare the crawl with XML sitemaps, Google Search Console landing pages, analytics landing pages, and any approved URL inventory. A URL found outside the crawl is an orphan candidate, not a confirmed orphan.

Verify that it is live, indexable, intended for users, and absent from internal links. Then decide whether to add contextual links, add it to an appropriate hub, consolidate it, redirect it, or remove it from supporting feeds.

P - Performance: Import PageSpeed Insights or other performance evidence when available, then review slow or unstable templates by business importance. A lab metric from one run should lead to template investigation and repeat testing, not a guaranteed ranking claim. Record whether the issue is page-specific or shared across a template.

Y - Your On-Page Signals: Only after crawlability, indexability, content purpose, architecture, and template behaviour are understood should you finalise title tags, meta descriptions, H1 and H2 structure, image alt text, and existing structured data.

Validate that each element reflects the page users actually see and does not conflict with canonical or indexation decisions.

Treat low content counts and duplicate flags as prompts for page-purpose review, not automatic rewrite instructions
Use crawl depth to investigate architecture and internal support; fix the linking path when the problem is structural
Verify orphan candidates against live status, indexability, purpose, and internal links before adding them back into the site
Separate sitewide navigation links from contextual links so template volume does not hide weak page-level support
Use performance imports for triage and repeat testing, while distinguishing a template problem from a single-URL problem
Finish with titles, descriptions, headings, alt text, and existing structured data only after upstream constraints are resolved

3Use Custom Extraction to Test Page Templates and Content Requirements

Custom extraction lets you collect a defined element from every crawled page, which is useful when the built-in columns do not answer the audit question. Open Configuration > Custom > Extraction and choose CSSPath, XPath, or regex.

Before running a full crawl, test each rule on several URLs from every relevant template. Record what an empty result means: missing content, a selector mismatch, content loaded after interaction, or a template where the field is not expected.

Use Case 1: Heading Hierarchy Audit Extract `h1` with a CSS selector and include the built-in H1 columns in the export. Review H1 text, H1 count, H1 visibility, and H1 template output. The goal is not to enforce an arbitrary identical pattern across every page; it is to confirm that the main heading is present where expected, describes the visible page, and is not duplicated by hidden mobile or desktop markup.

An empty extraction should be checked in the rendered HTML before it becomes a task. Confirm the H1 source before assigning the issue.

Use Case 2: Schema Markup Detection Use the XPath `//script[@type='application/ld+json']` to collect JSON-LD blocks. Presence alone does not establish validity, eligibility, or usefulness. Compare the extracted data with the visible page, check that the type matches the actual entity or content, and validate the implementation using appropriate testing tools.

Do not add markup only because another page type has it, and do not claim that FAQPage markup can produce a Google FAQ rich result.

Use Case 3: Word Count Approximation The built-in Word Count column can support review when its configuration matches the page content you intend to measure. If you also extract `//body`, be aware that navigation, footer, legal text, and repeated modules may inflate the count.

Use 300 words for informational content and 500 for commercial pages only as preserved review thresholds in this workflow, not as universal minimums. Compare suspected pages by template, intent, duplication, and visible main content before deciding on expansion.

Use Case 4: CTA Presence Check Extract a consistent call-to-action selector such as `.your-cta-class-name` when the business has defined where that element should appear. A missing result can indicate a selector issue, a conditional component, or a genuine omission. Verify the page manually and confirm that the call to action suits the page purpose before changing the template.

Custom extraction is most reliable when each rule has a written requirement, a sample of expected matches, a sample of expected non-matches, and a validation step. If the selector returns inconsistent results, stop and repair the rule rather than interpreting partial data as a site-wide finding.

Choose CSS selectors for stable element patterns and XPath for more specific document relationships; test both against real templates
An H1 extraction can be configured in under 10 minutes, but the result still needs rendered-page and template validation
JSON-LD extraction shows coverage and content, while a separate validation step is needed to assess syntax and page consistency
Use body text and Word Count as comparative signals, then review visible main content before recommending expansion or consolidation
CTA extraction verifies an agreed template requirement; it does not prove that the page converts or that the same CTA belongs everywhere
Save tested extraction rules with notes so future crawls use the same selectors and the same interpretation

4Combine Crawl Depth, URL Importance, and Search Evidence to Choose Internal-Linking Work

Crawl depth becomes decision-useful when it is combined with page purpose and observed search visibility. Export the crawl depth report from Reports > Crawl Analysis > Crawl Depth, then join it to the URL-level GSC data and Google Analytics data already imported or exported for the same property and date window.

Keep the source columns intact so anyone reviewing the sheet can distinguish crawl data, search data, analytics data, and manual business judgement.

The working matrix retains three dimensions:

- Crawl Depth (1-7+ levels from homepage) - Commercial Value (revenue attribution, conversion rate, or strategic priority) - Current Organic Visibility (GSC impressions and position)

Use the matrix to form hypotheses. A high-value URL with deep crawl depth and low visibility may be under-supported internally, but it may also have weak content, a canonical conflict, noindex directives, poor intent alignment, or limited demand.

Before adding links, inspect its status code, indexability, canonical target, inlinks, outlinks, sitemap presence, and query data. If those checks are sound, identify relevant source pages that naturally help a reader reach the target.

Do not rewrite the target solely because it is deep. First test whether the information architecture should expose it more directly. Add three to five contextually relevant internal links only where the references are useful and accurate.

A navigation or hub link may be appropriate when the page belongs in a durable hierarchy; otherwise, in-body links from related pages may be clearer. Avoid inserting links mechanically or repeating the same anchor text across unrelated sources.

Treat any observed change after implementation as an outcome to investigate, not proof that internal links alone caused it. Record the change date, source URLs, destination URL, anchor text, and other edits made at the same time.

Re-crawl to confirm that the new links are discoverable and that crawl depth or inlink counts changed as expected. Then review later GSC data using a comparable date window.

How to build the matrix in Screaming Frog::

  1. Run your crawl and connect GSC and GA4 APIs before crawling
  2. Go to Reports > Crawl Analysis and export the depth report
  3. Export the full URL list with GSC data (clicks, impressions, position)
  4. In your spreadsheet, create a column for crawl depth, one for impressions, one for commercial value (manual score or revenue data), and one for current average position
  5. Colour-code by quadrant: green (shallow + visible), amber (shallow + invisible or deep + visible), red (deep + invisible + commercially important)
  6. Prioritise every red URL for investigation before choosing an internal-linking, content, canonical, or consolidation task
Crawl depth is an observation about the tested path, so compare pages at depth six or deeper with pages at depth two or three without treating depth as a guaranteed crawl-frequency cause
Relevant internal links can improve discoverability and signal page relationships, but the source page, anchor, and user value should be documented
Use the matrix to distinguish an architecture hypothesis from content, canonical, indexability, or demand problems
Contextual internal links are often faster to implement than a full rewrite, but validate the target before adding them
A page at depth one or two with zero impressions needs broader diagnosis, including intent, indexation, canonical, content, and demand checks
Repeat the matrix quarterly with the same definitions because site structure and page priorities can change

5Audit Titles and Meta Descriptions With Query Evidence and Acceptance Tests

Screaming Frog is effective at finding missing, duplicate, long, short, and multiple title elements, but the flags do not tell you what the replacement should say. Start by exporting Page Titles with URL, indexability, canonical, main heading, pixel width, GSC queries, impressions, clicks, and average position where available.

Group findings by template so a system-level defect is not converted into dozens of manual edits, and treat the 60-character flag as a review signal.

The Title Tag Audit Process in Screaming Frog In the Page Titles tab, review Missing, Duplicate, Over 60 Characters, Under 200 Pixels, and Multiple. Compare the H1 where relevant. Missing titles require prompt investigation, but first confirm that the URL is indexable and intended as a search landing page.

For duplicates, sort by title and URL pattern. If 40 duplicate titles share a template, inspect the template variables, pagination, filters, and canonical rules before editing individual pages.

Do not treat 60 characters as a hard display limit. Screaming Frog may flag titles over 60 characters and also reports pixel width, which better reflects how a title may fit, although Google can still rewrite titles and display behaviour can vary.

A title with narrow characters at 65 characters can occupy less space than one with wide characters at 55 characters. Use width as a review signal, then prioritise clarity, intent, and accurate page description.

The CTR Optimisation Layer With GSC connected, filter pages with more than 200 impressions and a click-through rate below the site average. Then inspect the actual queries, positions, devices, countries, and date range.

A low rate may reflect a weak title or description, but it can also reflect ranking position, mixed intent, brand demand, SERP features, or irrelevant impressions. Create a snippet task only when the query evidence and page purpose support it.

Draft the new title from the page's real purpose and the language used in relevant queries. Keep the brand and differentiator only when they help the user choose the result. Do not promise an outcome the page cannot substantiate.

For validation, check that the deployed title is unique where uniqueness is intended, appears in the rendered HTML, matches the indexable canonical page, and does not exceed the agreed template rules.

A two to four weeks review window can be used as a first observation stage for recrawling and early GSC inspection, but it is not a guaranteed response time. Meta descriptions do not directly determine rankings; use them to summarise the page accurately and give a clear reason to visit.

If the result remains inconclusive, extend the comparison window, control for position and seasonality, and review whether Google is showing a different snippet.

Use pixel width and character count as review columns, while prioritising an accurate, useful title over a rigid length target
When duplicate titles share a URL pattern, repair the template or page model instead of editing each output manually
Use GSC filters to locate review candidates, then diagnose query relevance, position, device, and intent before rewriting
Align titles with the page's actual relevant query set rather than forcing an intended keyword that the page does not satisfy
Write meta descriptions as accurate result summaries that explain the page's value without unsupported promises
Investigate missing titles first on indexable pages that are intended to appear in search
Run the title review after importing search data when available so prioritisation is based on evidence rather than warning counts

7Check Redirects, Canonicals, and Response Codes Before Rewriting Page Content

A page cannot benefit from on-page improvements in the intended way when users and crawlers are sent elsewhere, the canonical consolidates signals to another URL, or the server response is unstable. Use Screaming Frog to map these conditions before creating content tasks.

For each priority URL, retain the original URL, final URL, response chain, indexability, canonical target, inlinks, sitemap status, and any search data.

Redirect Chains Open Reports > Redirect Chains and review URLs that pass through more than one redirect. The practical goal is a direct, intentional destination with no avoidable intermediate hop.

Before changing a chain, confirm that each source is obsolete, that the final destination is the closest relevant equivalent, and that no application or tracking requirement depends on an intermediate step. Consolidate eligible chains to a single 301 redirect and update internal links to the final URL.

Canonical Misconfigurations In the Canonicals tab, compare self-referencing, missing, multiple, and non-self canonicals. A non-self canonical is not automatically wrong; it may intentionally consolidate duplicate or variant pages.

Investigate when the canonical conflicts with internal links, sitemap inclusion, indexability, or the page's intended role. If page A points to page B, ensure that page B is the preferred equivalent and that users are not being asked to discover or convert on page A as though it were independently indexable.

Pagination such as page-2 and page-3 requires page-type review rather than a blanket canonical rule. Determine whether each URL contains unique, useful items, whether users and crawlers need to reach deeper items, and whether the implementation is consistent with the site's indexing plan.

Index Coverage via Response Codes Filter Response Codes for non-200 responses, then prioritise by internal links, sitemap presence, GSC impressions, and user importance. A 404 linked internally should be fixed at the source, redirected to a truly relevant replacement, restored, or intentionally removed from navigation.

A 301 chain containing 301 responses should be shortened where possible. A 5XX response on an important page requires prompt reliability investigation, but one crawl occurrence should be confirmed with logs or repeat checks before the issue is characterised.

For every 404 with historical GSC impressions, review whether the old content has a close live equivalent. Redirect only when the destination serves the same or a clearly equivalent need; otherwise, a 404 response can be more accurate.

Re-crawl two weeks after deployment as a later comparison stage, while also running an immediate targeted crawl to verify status codes and canonicals.

Review any redirect chain longer than two hops and collapse avoidable paths to a direct destination after confirming the intended mapping
A canonical pointing elsewhere requires intent and equivalence review; it does not automatically mean the source page is broken
Choose pagination behaviour from content uniqueness, navigation needs, and indexing intent rather than a single blanket rule
Fix internal links to 404 targets and redirect only when a genuinely relevant replacement exists
Treat 5XX responses on high-impression pages as urgent reliability signals and confirm them with repeat checks or server evidence
A self-referencing canonical is normally consistent for an indexable preferred URL, but it should still match the final live URL and page intent

8Turn the Crawl Into a Repeatable Monitoring and Validation Process

The audit is complete only when changes have been implemented, checked, and compared against a stable baseline. A later full crawl is useful, but it should not replace immediate validation of high-risk changes such as redirects, canonicals, robots directives, or template updates.

Define two stages: implementation verification soon after release, and trend review after enough comparable search or analytics data has accumulated.

Scheduled Crawl Cadence The paid version can schedule crawls and exports. A monthly crawl for small to medium sites under 5,000 pages and a bi-weekly crawl for large or fast-changing sites can be used as an operating practice, not as an official requirement.

Choose the cadence from publishing frequency, deployment risk, crawl cost, and team capacity. Save the same configuration, property selections, GA4 connection, and rendering mode for each comparable run.

Crawl Comparison Reports Use Reports > Crawl Comparison to load two crawl files and review additions, removals, status-code changes, title changes, canonical changes, and depth changes. Confirm that both crawls used compatible settings.

When a difference appears, inspect the underlying URLs and deployment history before labelling it a regression. A rendering or scope change can create hundreds of apparent differences without any page change.

Issue Velocity Tracking Maintain a tracking sheet for missing titles, broken internal links, redirect chains, pages at depth five or deeper, and orphan candidates. These counts are operational indicators, not a scoring system for rankings.

A rising orphan count may suggest that publishing and internal-linking workflows are misaligned, while a falling count may reflect consolidation, better links, or a changed discovery source. Add notes explaining major shifts.

The Quarterly Review Every quarter, repeat the full layered review across Content, Architecture, Navigation, Orphans, Performance, and Your On-Page Signals. Use the quarter as a planning interval, while assigning shorter validation windows to implementation checks and longer windows to search-performance interpretation.

Do not assume that every crawl difference requires a task. Close findings that are intentional, document exceptions, and carry unresolved items forward with the evidence still needed.

The finished system should produce a small, actionable queue: URL, issue, evidence, owner, action, acceptance test, validation crawl date, and result. When a result is inconclusive, state why and choose the next evidence source, such as rendered HTML, server logs, Search Console indexing reports, template code, or a manual page comparison.

Use scheduled crawls in the paid version when the cadence matches site change risk and team capacity
Crawl Comparison can expose regressions, but every difference should be checked against configuration and deployment history
Track issue counts over time as operational signals and annotate changes rather than presenting them as ranking predictors
Use quarterly full reviews for broad reprioritisation and shorter targeted crawls for implementation verification
Monthly checks of response codes, orphan candidates, and new title issues can catch fast-moving defects between broader reviews
Keep configuration consistent across comparable crawls; otherwise differences may reflect the crawler rather than the site

9What Most Guides Get Wrong

The first failure is starting with the issue tabs instead of the audit question. A useful crawl begins with a decision such as: which indexable pages are inaccessible from internal links, which templates create duplicate signals, or which high-impression URLs have weak snippets? Without that question, the export becomes an unranked inventory.

The second failure is treating every warning as an error. Screaming Frog reports observable conditions. It may show a long title, a non-indexable canonical, a redirect, or a low word count, but the correct action depends on page purpose, template behaviour, and search evidence.

The auditor must verify the rendered page, the source pattern, and the intended destination before recommending a change.

The third failure is changing pages without a baseline or acceptance criteria. If the initial crawl settings are not saved, the follow-up crawl may not be comparable. If the task says only 'fix title tags,' no one knows whether success means uniqueness, improved query alignment, correct templating, or simply removal of a warning.

Define the expected post-change state before implementation and record any cases where the crawl cannot resolve intent, business value, or indexation on its own.

10What Matters More Than Finding 50 Issues

Early Screaming Frog audits often become long lists because the software makes detection easy. A report with 400 flagged items can still be hard to use when it does not separate intentional conditions, template defects, page-specific problems, and unresolved questions.

The practical skill is to convert observations into decisions: identify the URL pattern, verify the rendered page, determine the intended search role, choose the smallest appropriate fix, and define the acceptance test.

The most reliable audits also distinguish implementation validation from performance interpretation. A targeted re-crawl can confirm that a title, redirect, canonical, or internal link was deployed. It cannot prove why rankings changed.

Search outcomes require comparable data, awareness of concurrent changes, and cautious language. A smaller queue with clear owners and tests is more useful than a second 400-item list built from warning counts alone.

11Your 30-Day Screaming Frog On-Page SEO Implementation and Validation Plan

Days 1-2

Define the audit question and scope. Configure the user agent, rendering, crawl rules, and authorised GSC/GA4 connections. Test sample templates, run the full crawl, and save both the configuration and baseline file.

Outcome: A reproducible baseline crawl with documented scope, settings, data sources, and known exclusions.

Days 3-4

Review the content layer. Extract H1s, compare the H1 field, existing structured data, and template-specific elements; compare Word Count and duplication findings with page purpose before classifying any page as thin or incomplete.

Outcome: A verified content worklist that separates missing elements, template defects, duplication, consolidation candidates, and intentional exceptions.

Days 5-7

Review architecture and navigation. Export crawl depth and All Inlinks, then join them to GSC and GA4 data. Investigate red-quadrant pages before deciding whether they need links, content, canonical changes, or consolidation.

Outcome: A prioritised internal-link and architecture queue with an evidence-based action and acceptance test for each URL.

Days 8-10

Compare crawled URLs with sitemap, GSC, analytics, and approved inventories. Verify each orphan candidate's live status, purpose, indexability, and current inlinks before assigning a link, consolidation, redirect, or removal task.

Outcome: A resolved orphan-candidate list in which intended pages are discoverable and obsolete or duplicate URLs have a documented disposition.

Days 11-14

Import PageSpeed Insights data where available. Group poor Core Web Vitals observations by template, repeat representative tests, and send validated template or page defects to the development team with reproduction evidence.

Outcome: A performance backlog ranked by page importance and evidence quality rather than by a single isolated score.

Days 15-20

Audit missing, duplicate, and over-width titles; meta descriptions; H1/H2 structure; and image alt text. Use GSC query and CTR data to select candidates, then draft changes that match the visible page and intended canonical.

Outcome: A controlled rewrite queue with deployment rules, query evidence, and post-release checks for each page or template.

Days 21-25

Use the All Inlinks export to inspect source pages, destination pages, anchors, link position, and target status. Add or revise contextual links only where they improve navigation and accurately describe the destination.

Outcome: A cleaner internal-link map in which priority pages receive useful, verifiable support from relevant sources.

Days 26-28

Inspect redirect chains, canonical targets, and response codes. Shorten eligible chains, correct conflicting canonicals, update broken internal links, and redirect 404s only when a genuinely relevant destination exists.

Outcome: Direct crawl paths, consistent canonical decisions, and documented treatment of broken or retired URLs.

Days 29-30

Run targeted validation crawls, then schedule monthly comparisons with the saved configuration. Create the issue tracking sheet and document unresolved findings, the evidence still needed, and the next quarterly review scope.

Outcome: An ongoing monitoring process with a baseline, acceptance tests, implementation checks, and a clear path for inconclusive findings.

Define the audit question and scope. Configure the user agent, rendering, crawl rules, and authorised GSC/GA4 connections. Test sample templates, run the full crawl, and save both the configuration and baseline file.
Review the content layer. Extract H1s, compare the H1 field, existing structured data, and template-specific elements; compare Word Count and duplication findings with page purpose before classifying any page as thin or incomplete.
Review architecture and navigation. Export crawl depth and All Inlinks, then join them to GSC and GA4 data. Investigate red-quadrant pages before deciding whether they need links, content, canonical changes, or consolidation.
Compare crawled URLs with sitemap, GSC, analytics, and approved inventories. Verify each orphan candidate's live status, purpose, indexability, and current inlinks before assigning a link, consolidation, redirect, or removal task.
Import PageSpeed Insights data where available. Group poor Core Web Vitals observations by template, repeat representative tests, and send validated template or page defects to the development team with reproduction evidence.
Audit missing, duplicate, and over-width titles; meta descriptions; H1/H2 structure; and image alt text. Use GSC query and CTR data to select candidates, then draft changes that match the visible page and intended canonical.
Use the All Inlinks export to inspect source pages, destination pages, anchors, link position, and target status. Add or revise contextual links only where they improve navigation and accurately describe the destination.
Inspect redirect chains, canonical targets, and response codes. Shorten eligible chains, correct conflicting canonicals, update broken internal links, and redirect 404s only when a genuinely relevant destination exists.
Run targeted validation crawls, then schedule monthly comparisons with the saved configuration. Create the issue tracking sheet and document unresolved findings, the evidence still needed, and the next quarterly review scope.

Frequently Asked Questions

How many pages can I crawl with Screaming Frog for free?

The free version allows up to 500 URLs in a crawl. That can cover a small site, a subfolder, or a representative template sample. The paid licence removes the crawl limit and includes features such as scheduled crawls, crawl comparison, API integrations, and Link Opportunities.

Choose the version from site size and the evidence you need. For a site larger than the free limit, you can still use a scoped sample, but document that excluded URLs may contain additional issues.

What should I audit first in Screaming Frog?

Start with the audit question, crawl scope, and configuration. Then review content purpose, architecture and crawl depth, navigation and internal links, orphan candidates, performance evidence, and finally titles, descriptions, headings, alt text, and existing structured data.

This order prevents you from rewriting a page before confirming that it is crawlable, indexable, canonicalised correctly, and supported by the site's internal structure.

Does Screaming Frog work for JavaScript-rendered websites?

Yes. In Configuration > Spider > Rendering, select 'JavaScript' and test representative pages before the full crawl. Compare the rendered HTML with the raw response and confirm that main content, H1 tags, canonicals, and links are present.

JavaScript rendering uses more resources and can expose content that is absent from the initial HTML. When results differ, record the rendering mode and use the crawl that matches the audit question.

How do I find orphaned pages in Screaming Frog?

Enable the relevant XML sitemap source and compare sitemap URLs with URLs discovered through internal links. Also compare GSC and analytics landing pages when authorised. A URL that appears in those sources but not in the crawl is an orphan candidate.

Verify that it returns a live status, is intended for users, is indexable, and truly has zero inlinks. Then add a useful contextual link, place it in an appropriate hub, consolidate it, redirect it, or remove it from supporting feeds according to its purpose.

Can Screaming Frog help me improve click-through rates, not just rankings?

Yes. With the Google Search Console API connected, you can compare page-level impressions, clicks, CTR, position, and query data with titles and meta descriptions. Filter high-impression, below-average CTR pages as review candidates, then check query relevance, ranking position, device, country, and the displayed snippet.

Rewrite only when the current snippet does not accurately or clearly present the page. Re-crawl to verify deployment and use a comparable later window to assess the result.

How often should I run a Screaming Frog crawl?

For a site under 5,000 pages, a monthly full crawl and a quarterly broader review can be a workable operating cadence. A fast-changing or high-risk site may justify bi-weekly checks. The correct frequency depends on publishing, deployments, crawl cost, and team capacity.

Include new 404 findings in the targeted review. Keep the same configuration for comparable crawls, and use targeted crawls immediately after important redirect, canonical, template, or robots changes.

What is the difference between canonical tags and redirects in Screaming Frog?

A 301 redirect sends users and crawlers to another URL, while a canonical leaves the source accessible and declares a preferred URL for consolidation. In Screaming Frog, review redirects in Response Codes and Redirect Chains, and review canonical declarations in Canonicals.

Neither should be changed from the report alone. Confirm the intended preferred page, content equivalence, internal links, sitemap entries, and final status before implementing a redirect or replacing a canonical.

THIRTY SECONDS TO START

You've read enough.Your own data says more.

Connect your site and see it yourself: your rankings, your gaps, your blockers, and what AI tells your buyers. The plan and the priced options follow within 36 hours.

Your access code by SMS. We never call.No payment