Skip to content
    Search & the Web

    Who pays for the crawlable web

    Crawler traffic is now a material infrastructure cost for small sites, and the tools for managing it are blunt.

    By Priya Raman6 min read
    Fibre optic strands glowing against a dark background
    Fibre optic strands glowing against a dark background

    Site operators increasingly report that a majority of their requests come from automated clients. For a large publisher that is an accounting line. For an independent site it can be the difference between a hosting bill of tens and hundreds.

    The controls available

    • robots.txt, which is a request rather than an enforcement mechanism and is honoured inconsistently.
    • Rate limiting at the edge, which works but risks blocking the crawlers that still send audiences.
    • Licensing arrangements, which are practical for large publishers and largely unavailable to small ones.

    The asymmetry

    A crawl is cheap to initiate and comparatively expensive to serve. Nothing in the current protocol layer corrects that imbalance, and the sites least able to absorb the cost are the ones with the least negotiating leverage.

    Proposals for machine-readable licensing terms exist. Adoption, so far, follows the size of the publisher.

    About the author

    Priya Raman

    Security & Web Correspondent

    Priya Raman reports on authentication, software supply chains and the changing shape of search and the open web. Her work focuses on how security decisions affect ordinary users.

    Related stories