SEO Recap covering August 10, 2026: Clarity branded citations and Common Crawl visibility tools
Daily summary of what matters in SEO, GEO, AEO, and AI search generated with Claude Code (beware of hallucinations)
Generative Engine Optimization (GEO) & Answer Engine Optimization (AEO)
Microsoft Clarity Adds Branded and Non-Branded AI Citation Labels
Source: Search Engine Journal
- Clarity's August 3 update added branded labels, filters, citation counts, and Share of Authority reporting to its AI Citations dashboard, letting marketers separate whether an AI system searched for their brand by name or found their content while researching a broader topic.
- Microsoft defines Share of Authority as the percentage of citations attributed to your domain versus other cited domains, calculated daily across query-days where your domain was cited. It reflects visibility only within those included queries, not the whole AI search market.
- Clarity tracks citations separately from AI referral traffic. A page can earn citations without any visits, so teams should measure three distinct stages:
- Citation visibility (how often AI references your pages)
- AI referral traffic (sessions arriving from AI assistants)
- Business results (leads, sales, revenue from those sessions)
- Grounding queries are the AI system's behind-the-scenes searches and may differ from the user's prompt, so they show AI retrieval behavior rather than customer intent or buying stage.
- Microsoft calls the dashboard a representative view across supported AI experiences, not a complete record. Very low-volume activity may be excluded, totals can vary between page and query views, and brand classification has edge cases for common-word or multi-brand names.
A Free Tool to Audit Your Visibility in Common Crawl's AI Training Data
Source: Search Engine Journal
- Common Crawl fetches over 2 billion pages monthly (2.14 billion in the July 2026 crawl) and its archive feeds most AI training sets. Mozilla found 64% of LLMs released between 2019 and 2023 trained on Common Crawl data, and GPT-3 was ~60% filtered Common Crawl by training weight.
- HTTP Archive data cited by Chris Green shows ~492,000 sites name CCBot in robots.txt, ~95% to block it. Since July 1 2025 Cloudflare blocks AI crawlers by default on new domains, and its managed robots.txt reached 3.8 million domains, so many blocks are platform defaults rather than deliberate choices.
- The Common Crawl Visibility Checker automates Common Crawl's 18-page manual audit, returning four panels:
- Captures per crawl across the last 12 monthly crawls
- Robots.txt history with block start dates
- Whose block template it matches (Cloudflare, Squarespace, plugins)
- A live probe fetching the homepage as CCBot against a browser to catch edge/WAF challenges
- To improve inclusion: fix CDN/edge challenges first, allow CCBot, add a sitemap to robots.txt, render on the server (CCBot runs no JavaScript, and the stored limit rose from 1 MiB to 5 MiB in March 2025), and earn links from well-connected domains since crawl priority follows harmonic centrality.
- Being in Common Crawl does not guarantee inclusion in a model, since every training set filters the crawl and frontier labs stopped disclosing their data mixes around 2023.
AI in Search / AI Overviews
Meta Reportedly Crawling the Web to Build Its Own Search Engine
Source: Search Engine Roundtable
- Pieter Levels posted screenshots on X showing Meta crawling his properties, and reported that Meta is allegedly building its own web index so its AI's web searches do not rely on Google, which could otherwise use the queries for training.
- Facebook previously partnered with Bing to power web search features before dropping the partnership, so an in-house index would reduce dependence on third-party partners for AI.
Google Tests New AI Buttons on the Desktop Home Page
Source: Search Engine Roundtable
- Google is testing Create images, Ask about files, and Brainstorm buttons below the search box on desktop, alongside the existing AI Mode button, plus sign, Lens, and Voice options, replacing the classic Search and I'm Feeling Lucky buttons for some users.
- VP of Product Robby Stein confirmed it as a small desktop test to help people find new things they can do with Search, with no impact on how the core search box works.
Technical SEO
Google's Illyes Says Hreflang URLs Are Not Indexed in the Proper Sense
Source: Search Engine Roundtable
- Gary Illyes said hreflang alternates become alternate names stored against the canonical URL rather than being indexed in the proper sense, so the specific URL is kept as an alternative to the main page.
- Alternate names can still appear in results when a query deserves one, with site: queries the most prominent case. The alternate URL maps to the canonical page that is actually indexed.
Published on