A Technical SEO Audit Checklist for Small Sites

Your workspace during a site audit

Small sites get audited badly. Most checklists were written for enterprise platforms with thousands of templates, so they front-load log file analysis and crawl budget math that a forty page site will never need. Meanwhile the issues that actually suppress a small site tend to be dull: a page nobody links to, a stray noindex, a navigation structure that buries the money pages three clicks deep for no reason.

What follows is a working order of operations for auditing a small site, plus a note at each stage about whether that fundamental also affects retrieval by AI assistants and answer engines. The short version is that most of it does, because retrieval systems still depend on crawlers, and a URL that cannot be fetched and parsed cannot be quoted back to a user.

Start with crawlability, because nothing downstream matters without it

Begin by confirming that a machine requesting your pages gets a clean response. Check the robots.txt file first and read it line by line rather than assuming it is fine. Disallow rules written during a staging phase have a habit of surviving launch, and a blanket disallow on a directory can quietly remove an entire section from consideration.

Then crawl the site yourself with any crawler you trust and look at status codes in aggregate. You want to know how many URLs return 200, how many redirect, how many chain through multiple redirects before resolving, and how many return 404 or 500. On a small site you can often eyeball the whole list. Pay attention to internal links pointing at redirects, because those are cheap to fix and they clarify the crawl graph.

Next, check whether your primary content is present in the initial HTML response or whether it is assembled client side by JavaScript. View the raw source, not the rendered DOM in developer tools. If the body copy only appears after scripts execute, you are relying on the crawler to render, and rendering is not guaranteed across every system that might want your content. This is the single crawlability issue with the biggest gap between traditional search and AI retrieval. Established search crawlers invest heavily in rendering. Many of the fetchers used by AI assistants and answer engines are far more conservative, and content that requires execution to appear is content they may simply never see.

Also worth checking: server response times under normal conditions, whether your pages are gated by cookie walls or interstitials, and whether any aggressive bot protection is returning 403 responses to non-browser user agents. Blocking scrapers and blocking retrievers can be the same action, so decide deliberately rather than by default.

Indexation: confirm what is eligible and what is duplicated

Crawlable is not the same as indexable. Work through the signals that control eligibility one at a time.

Look for meta robots noindex tags and X-Robots-Tag headers across the site. Headers are the sneaky ones because they are invisible in the page source. Then examine canonical tags: every indexable page should canonicalize to itself unless you have a specific reason otherwise, and cross-domain or all-pages-to-homepage canonicals are usually accidents.

List your duplicate and near-duplicate surfaces. On small sites these commonly come from parameter variants, trailing slash inconsistencies, http and https both resolving, www and non-www both resolving, pagination, and tag or category archives that contain only excerpts. Pick one canonical form per URL, redirect the rest, and make sure your internal links use the canonical form rather than relying on redirects to sort it out.

Then audit for thin pages. A small site does not need many URLs, and a handful of substantial pages will nearly always outperform a wide spread of shallow ones. If a page exists only to target a keyword variant and says nothing a reader could not get elsewhere on the site, consolidate it.

Finally, generate or review your XML sitemap and confirm it contains exactly the canonical, indexable, 200-returning URLs and nothing else. Sitemaps full of redirects and noindexed pages are a signal of neglect.

Indexation work pays off in AI retrieval too, though indirectly. Retrieval systems that lean on their own indexes benefit from a site with one clear URL per topic, and systems that lean on a search index inherit whatever duplication confusion you left behind. Consolidating near-duplicates also makes it likelier that a single strong page is the one retrieved and cited, rather than attention splitting across three weak variants.

Internal linking and page architecture on a small footprint

With a small site the architecture question is not depth management at scale, it is whether every important page has enough internal support to look important.

Map your click depth from the homepage. On a site under a few hundred URLs, almost nothing should sit more than three clicks from the entry point. Then flip the view and count internal inbound links per page. Orphan pages, reachable only via the sitemap or a direct URL, are the classic small-site failure. So are pages with a single inbound link from a footer.

Check your anchor text for descriptiveness. Links reading "learn more" and "click here" transmit nothing about the destination. Descriptive anchors help search engines understand the target page and they help language models understand the relationship between your pages when they process them.

Think about architecture in terms of topic grouping rather than menu aesthetics. Related pages should link to each other, and a hub page that explains a subject area and links out to its supporting pages does genuine work. It concentrates relevance signals, it gives readers an obvious next step, and it gives a retrieval system a page that summarizes a topic with pointers to detail.

For AI answer surfaces, architecture matters in a slightly different way. Retrieval usually operates at the passage level, which means the unit of competition is often a section rather than a whole page. That rewards pages with honest heading hierarchies, one idea per section, and self-contained passages that make sense when lifted out of context. If a paragraph only parses because of a sentence three screens above it, it travels poorly.

On-page structure that serves both readers and machines

Once crawling, indexation and linking are sound, review page level structure. Each page should have a unique title that describes it accurately, a single H1 that matches the page's actual subject, and a heading hierarchy that descends in order without skipping levels for styling reasons.

Add structured data where it genuinely describes the page: organization details, article metadata, product or service information, frequently asked questions where they truly exist. Structured data does not create authority, but it removes ambiguity about what a page is and who published it.

Confirm the basics of accessibility and semantics, since they overlap heavily with machine readability. Real heading tags rather than styled divs, alt text on informative images, tables marked up as tables, lists marked up as lists. Content buried in an image or a PDF with no text layer is content that cannot be retrieved or quoted.

One more habit worth building: write the first two or three sentences of each section so they answer the section's implicit question directly. That is good editing regardless of audience, and it also means that when a passage is retrieved in isolation, the useful part arrives first.

Working the checklist in order

The sequence matters more than the individual items. Fix fetchability before indexation, indexation before linking, linking before on-page polish. Auditing out of order produces a long list of fixes where half of them are invisible because a robots directive upstream is suppressing the pages they live on.

Run the crawl again after each round of changes and compare, rather than assuming the fix landed. Small sites have the advantage here, since a full recrawl takes minutes and you can inspect the whole URL set by hand.

And resist treating AI visibility as a separate project with a separate checklist. The overlap is large. Server-rendered content, clean status codes, consolidated URLs, descriptive internal links, honest heading structure and self-contained passages serve classic search and AI retrieval at the same time. The genuinely AI-specific work, the part about how answers get assembled and which sources get named, sits on top of these fundamentals. It does not substitute for them. More on that side of the practice in Generative Engine Optimization.