Skip to main content
Web Development

What a Technical SEO Audit Actually Covers

a laptop and a monitor showing charts and performance reports

Most technical audits I have reviewed were lists. Four hundred rows, a severity column, a red traffic light on rows with the worst numbers. Rows are easy to produce and easy to ignore, because nobody can tell from a list which item is causing the loss and which is merely irritating. A useful audit is closer to a diagnosis: a small number of specific causes, evidence for each, and an order in which to treat them.

What follows is roughly what a competent technical audit covers, in the order it gets done, including the parts that are frequently skipped because they are tedious. It is written for the person who has to decide whether to act on the findings, not for the person running the crawler.

Establishing what the site actually is

Before crawling anything, you want a picture of the architecture. Not a diagram — an inventory. How many distinct templates are there, and which pages use each one? Is content stored in the database, in flat files, or in a repository alongside the code? Are there subdomains serving different purposes, such as a staging copy, a legacy section, or an internal tool that happens to be publicly resolvable?

printed documents, a clipboard and a pen spread across a desk

This step changes the character of everything after it. A site with eleven templates and clean routing has a very different problem profile from one with four hundred URLs generated by a plugin with two hundred query-string parameters. Both will produce a large crawl. Only one of them is worth auditing template by template.

Ask for the CMS version, the plugin list, the server software, and whether anyone can edit templates without engineering help. That last question determines whether a fix ships in an afternoon or waits three weeks. An audit that produces findings nobody is able to act on is a document, not a plan.

At SmartEdge IT Solutions this stage is usually the shortest part of the engagement and the part clients are most surprised by. The architecture questions sound like project management until you realise that the answer to most of them changes what the crawl is even looking at.

What the crawl tells you, and where it stops

The crawl is the widest and least expensive source of evidence. It gives you every URL the site exposes to a crawler, the response code, the redirect chain, the canonical tag, the title, the meta description, the heading structure, the word count, and the internal link graph.

Used well, the internal link graph is the most under-rated output of the whole exercise. You can see which pages have authority pointed at them, which important pages are linked from nowhere, which sections are orphaned by a change to the navigation three months ago, and where a crawl budget is being spent. None of that appears in a rank tracker, and all of it shapes whether the pages you care about get discovered promptly.

Where the crawl stops is just as important. It does not see pages blocked in robots.txt, so a crawl will cheerfully report a clean indexation profile for a section you have excluded for years. It does not see the pages your users reach through a POST request, an in-app search, or a parameter combination. It does not tell you whether a page that returns a success code actually contains the content, only that something came back. Each of those blind spots needs another source, which is the subject of the next few sections.

Indexation: separating pages by what should happen to them

Indexation problems are best approached as a classification exercise, not a fix list. You take every URL the site can generate and assign each one to one of a small number of states: this must be indexed, this must not be indexed, this should be indexed but is not, and this should not exist at all.

printed documents, a clipboard and a pen spread across a desk

That last category is the one people resist creating. Parameter combinations, paginated views beyond a certain depth, search results within the site, print views, session identifiers — these URLs are frequently reachable, frequently linked from somewhere, and frequently indexed. Each one that gets indexed dilutes the signals for the pages you actually want found, and each one that gets crawled costs time.

The mechanisms are unglamorous and effective: canonical tags pointing at the version you want, noindex where you are confident the crawler can read the meta tag, robots rules for crawl control rather than index control, and parameter handling on the server so the duplicates are not generated in the first place. Be careful with the distinction between disallowing a URL and removing it — a blocked URL cannot be de-indexed, because the instruction to de-index has to be read first.

Canonical tags are the most abused tool in this area. A canonical pointing to a page that is itself blocked, redirected, or soft-404 is worse than no canonical at all, because it consolidates signals onto a destination you have already told the engine is not worth having.

Rendered content versus source content

Any site that assembles pages in the browser needs a specific kind of check, and this is where a surprising number of sites discover a serious problem: the content is in the JavaScript bundle, not in the served HTML. The crawler can execute JavaScript, but not always quickly, not always completely, and not always for every bot. On a large site, “sometimes, eventually, on a good day” is a materially different proposition from “present in the initial response”.

a close view of a screen showing an application's interface

The test is unglamorous. Fetch a sample of pages and look at what is actually in the response body before any script runs, then compare it to what a browser renders. Where the two diverge, you have found either a rendering dependency you should remove or a rendering cost you have to accept knowingly. Our React development and Angular development work sits on top of server-side rendering for exactly this reason, and the trade-off is worth understanding before a framework is chosen rather than after.

Two related things get checked here. Metadata must exist in the source, not be injected client-side, because titles and structured data are read from the initial response. And the content of the primary content area should not depend on interaction — a category that only appears after a click may never be discovered at all.

Structured data, and what it is not for

Structured data gets more attention per unit of real impact than anything else on this list. It is worth auditing, but the framing needs correcting, because structured data is a machine-readable description of something already on the page. It does not create visibility.

What to check: whether the types used match what the page actually contains, whether required properties are present, whether anything claims content a visitor cannot see, and whether the same entity is described differently on different pages. That last one matters more than people expect, because inconsistent descriptions make it harder to associate the pages with each other, and association is what you are describing in the first place.

What not to do: treat a validation warning as a priority on its own. Warnings are often about fields that do not affect eligibility for anything. And do not add markup to pages that are thin, because a description of an empty page is still a description of an empty page.

Log files settle what crawls cannot

If the site has an application or edge layer that logs requests, those logs are the most honest dataset available, and they are routinely ignored. They show what actually happened, including traffic from crawlers you never identified in a crawler list, response codes the crawl never triggered because it did not request that URL, and the timing distribution that reveals a backend struggling under specific query patterns.

a close view of a screen showing an application's interface

Two questions usually need them. What is being crawled that is not in the site map or the internal links — which often turns out to be a feed, a REST endpoint, or an old query string still accepted by the application. And how often are important pages crawled, compared with how often unimportant ones are. When the answer is “the paginated archive gets crawled four times a day and the product page once a fortnight”, you have found the actual problem, and no amount of on-page tweaking will fix it.

Getting at logs properly involves the infrastructure side of the estate rather than the front end, and it is worth treating as a small piece of DevOps work in its own right rather than something to bolt on at the end.

Performance findings: separating signal from noise

Performance belongs in a technical audit because it changes how pages are indexed and served, not because the audit should become a synthetic benchmark report. The useful part is a short list of causes, each with an estimated effect on the templates where it occurs.

a notebook open on a desk with a pen, a phone and printed sheets beside it

Server response time under realistic load, cache behaviour and revalidation, image weight and format, font loading strategy, the amount of JavaScript shipped to a page that does not need it, and the render-blocking pattern of the critical path. Each of these has a limited number of fixes and a reasonably predictable outcome. Third-party scripts deserve their own line because they are usually the largest single avoidable cost, and because fixing them requires a commercial conversation with whoever placed them.

Where it is easy to mislead: a single lab run on a developer laptop tells you very little, and a long list of automated “opportunities” mostly measures the tool. I would rather see a table of templates, each with its actual response times from field data and a named bottleneck, than three hundred rows generated by a scanner. Our full website audit is built that way deliberately, and the performance reasoning connects to what we cover separately on server monitoring and backup.

Turning findings into a queue

Sorting is where most audits fail to deliver anything. A finding becomes a task only when it has an owner, an estimate, and a place in a sequence. I sort on two axes: how much of the site it affects, and how confident anyone is that fixing it will change anything.

Template-level defects go first, because they apply to everything using that template and one fix removes hundreds of instances. Site-wide infrastructure issues come next — response times, crawl waste from duplicates, internal linking for orphaned sections. Genuinely one-off content problems come last, and often end up as a content brief rather than a code change.

Then there is the honesty question. Some findings will be listed as “no action”. A redirect to a page that is genuinely useful, a low-traffic template with a mediocre title, a slow third-party widget nobody can remove — these are real observations with no cost-effective response. Saying so is more useful than padding the plan with work nobody will do.

A queue that SmartEdge IT Solutions hands over is written so that a developer could pick it up without attending the audit call, and so that whoever signed it off could tell which items were deliberately deferred. That second column matters more than it sounds: it is the difference between a plan and a wish list.

What a technical audit deliberately leaves out

Content quality, keyword strategy, backlink profiles and competitor analysis are not technical. They belong in a broader search engine optimisation engagement, and an audit that blurs the boundary usually spends most of its length restating that content is thin without explaining what to write instead.

a close view of a desk with a keyboard, a notebook, a pen and a coffee cup

A single technical audit is also a snapshot. Crawl behaviour, indexation decisions and rendering change over months, so the more useful arrangement for most sites is a baseline once, then periodic re-checks once something is being fixed. How we run that engagement is described on our process page, including what is included and what falls outside the scope.

If you want the audit itself, start there rather than with a conversation about retainers. Tell SmartEdge IT Solutions what the site is and what has changed recently, and the useful conversation about scope will happen faster.

Editorial profile

Olivia Bennett Search and Answer Visibility Editor

Olivia Bennett writes for SmartEdge IT Solutions about how people find a service, from technical SEO and structured content through to answer engines and generative search. She reads search console reports for a living and treats an unreachable page as an unfinished one.

Also 2 articles in the Insights archive.

← Back to Blog