SEO Recap covering October 5, 2026: Google's crawl timing data and a looming core update

Built by Stephanie Chung·6 min read·8 stories

Daily summary of what matters in SEO, GEO, AEO, and AI search generated with Claude Code (beware of hallucinations)


Generative Engine Optimization (GEO) & Answer Engine Optimization (AEO)

Mueller on AI crawlers, sitemaps, RSS, and llms.txt
Source: Search Engine Journal

  • John Mueller says AI training crawlers usually offer no way to submit a sitemap, so sites that want to be found should stick to the generic sitemap.xml name or focus on RSS feeds, which are easier to discover from a page's HTML head.
  • Mueller said he has seen AI crawlers access his sitemap and RSS files in his own server logs, though he could not name them or say what they do with the files.
  • llms.txt cannot replace an XML sitemap. Google's systems treat the Markdown file like an HTML sitemap and cannot use it as a sitemap because it lacks the strict format.
  • A valid public sitemap showing "Couldn't fetch" usually comes down to host load or crawl demand, and crawl demand is often based on perceived site quality rather than anything in the file.

Turning an AI brand mention into action (Ahrefs webinar)
Source: Search Engine Journal

  • Ahrefs' Constance Tan recommends organizing AI monitoring around 3 question types:
    • Problems customers need to solve, capturing early discovery
    • Positioning and comparison questions that reveal which brands AI recommends
    • Factual questions testing pricing, capabilities, and availability
  • Fix owned pricing and product pages first, then pursue third-party corrections you can realistically secure. One Ahrefs outreach contacted 26 authors, 10 replied, and 4 updated their content.
  • Before treating a missing citation as a content gap, check bot access for firewall restrictions, broken URLs, timeouts, and pages that rely on JavaScript to display key information.
  • Review AI visibility alongside conversions and customer attribution rather than treating a citation as a sale.

Technical SEO

Gary Illyes shares Google's crawl, index, and serve timings
Source: Search Engine Roundtable

  • Gary Illyes presented internal Google data on how long crawling, indexing, and serving take at the Search Central Live Deep Dive in Barcelona, with typical and slowest figures for each process.
  • Crawling benchmarks:
    • New URL discovery ~20 hours (weeks to never)
    • Sitemap processing ~24 hours (up to 14 days or never on quality)
    • Refresh of a known URL ~30 days
  • Indexing end to end runs ~1.5 hours typically but months or never on quality issues, with rendering taking seconds to render and hours in the queue.
  • Serving: core update recovery takes 3-6 months (up to a year), while spam update changes land in 1-2 weeks.
  • Illyes cautioned the numbers were an exercise to see if the audience could relate to internally pulled figures, and that linked processes make delays stack.

Google's hidden product data layer in ecommerce
Source: Search Engine Journal

  • Google builds a persistent memory of your products from crawls, feed history, and third-party sources, stitched together and surfaced in ways invisible to your feed and schema audits.
  • 3 price features can conflict with reality:
    • Sale price badges require a 5-90% discount with the base price submitted for 30 of the past 200 days in the UK
    • Price drop badges are auto-generated from Google's 60-day average, with a "Was" figure you never submit
    • Price drop rich snippets pull the reference price from Google's historical records, not your markup
  • Google maintains its own image index and keeps old product image URLs unless they return a 404. Shopify, Magento, and WooCommerce commonly leave retired images live on the CDN.
  • Cross-platform data exchange is suspected, where a Merchant Center feed error can trigger simultaneous rejection in Google Ads and Amazon.

Organic Search & Algorithm Updates

Google explains its 2026 spam update surge and AI detection
Source: Search Engine Roundtable

  • Google says it runs more spam updates because there is far more new content than before, and it is now using AI to catch more of it.
  • Gary Illyes said Google filters out 40 billion spam pages a day, a figure he noted is not new.
  • 3 reasons Google keeps updating how it judges quality:
    • Content formats have expanded beyond 10 blue links, each needing its own quality judgment
    • Content breadth has grown, so core updates target pages not domains
    • Content issues like cloaking, doorways, scraped content, and link spam
  • Illyes warned that scaled, low-effort content is now becoming a bigger problem than link spam, and AI slop is increasing in the Discover feed.

What to expect from Google's next core update
Source: Search Engine Journal

  • Google is in the midst of a 2-week spam update, which may be clearing the table before a holiday-season core update.
  • Updated helpful content guidance stresses main content quality and user satisfaction, and defines high-quality content as fulfilling its purpose and reflecting meaningful effort, originality, skill, and accuracy.
  • Google's definition of main content now extends beyond text:
    • Interactive features like calculators, online tools, and games
    • User-generated contributions such as forum posts and reviews
    • Tabbed or expanded content
    • Titles and headings that summarize the page topic
    • Primary text and media
  • New guidance warns against deceptive authorship (AI-generated headshots, made-up names, false credentials) and urges manual fact-checking of all AI-generated content before publishing.

AI agent scraping threatens rank tracking
Source: Search Engine Journal

  • Cloudflare reports AI agent traffic grew more than 1,700% over the past year, with its network handling ~115 million HTTP requests a second (up from 63 million at the end of 2024) and peaks above 150 million. More than half of internet traffic this year was not human.
  • Ryan Jones of SERPrecon says rank trackers are failing because Google is blocking AI trackers and AI itself, and the trackers get caught up in it.
  • AI Mode queries are computationally more expensive to serve, and Google reduced AI Mode response cost to its lowest level since launch through hardware and engineering optimizations.
  • Google's adversarial anti-scraping stance may be aimed at stopping companies from reverse engineering its results through distillation, not just at commercial rank trackers.

OpenAI brings invisible text watermarks to ChatGPT in the EU
Source: Search Engine Journal

  • OpenAI will add an invisible watermark to qualifying ChatGPT and Codex text for EU users across all plans over the coming weeks, with no global default rollout.
  • Its textGrain method, at a 1% false positive target, found the watermark in ~80% of 200-token passages and ~95% of 400-token passages for flexible content like psychology, and much less for content like math.
  • Editing defeats it. Replacing 10% of words with synonyms in 400-token passages dropped detection from ~92% to 66%, and 25% replacement dropped it to 17%.
  • API users worldwide can opt in to watermarking (off by default), but the text detection tool is limited to approved researchers, not the public.
  • The rollout answers EU AI Act Article 50 transparency rules, with pre-market AI systems given until December 2 to meet marking and detection obligations.
Published on