Crawling

Three panel diagram of log file analysis for seo: raw googlebot access log lines classified as keep or waste, a four step verify and segment workflow, and a crawl budget verdict bar where only 44 percent of hits reach indexable canonical 200 urls.

Log File Analysis

Log file analysis is the practice of examining a server's raw access logs to see exactly…
Learn More
Annotated robots. Txt example showing user-agent, disallow, allow, and sitemap directives next to the difference between crawl control and index control.

Complete Guide to Robots.txt Configuration

Comprehensive guide to robots.txt covering implementation, best practices, troubleshooting, and advanced techniques.
Learn More
Diagram of cloaking: a server checks the user-agent and ip, serves googlebot a clean keyword-rich page, serves humans a different page, and google detects the mismatch as a spam-policy violation.

Disallow Humans (Bad Cloaking)

A cloaking case study seeing if Google penalizes the website if it were intentionally decieved.
Learn More
Line chart showing the twitter sistrix visibility index falling 32 percent within 24 hours after crawlers met a login wall, with before and after crawl states.

Twitter down 32% in 24 hours - SISTRIX

A case study from Sistrix on the impact of Twitter blocking Google from crawling and indexing.
Learn More
Diagram of how search engine crawling discovers and fetches pages before indexing on seoprocheck. Com

Crawling

Crawling is the process by which search engines discover and download pages on the web using…
Learn More
Two lane flow diagram of google two wave indexing, showing html indexed immediately while javascript waits in the render queue before rendering and indexing, with a bar comparison of about 9 times more time for javascript.

Rendering Queue: Google Needs 9X More Time To Crawl JS Than HTML

An experiment to find out how long it would take Googlebot to crawl JavaScript vs HTML…
Learn More

Get new blog posts by email: