For publishers: our crawler
This page is the information page the ARCEL crawler's User-Agent points to. The crawler reads the news feeds of the publishers in ARCEL's reviewed source list and, for sources whose licence allows full text, the article pages those feeds link to. Editors then review every story before readers see it, and readers always get a link to the publisher's original article.
Identity
The crawler identifies itself as ARCEL-DailyDose-Bot/<version> (+<this page>), adding a contact address once one is configured. robots.txt rules naming ARCEL-DailyDose-Bot apply to it.
How it behaves
- Pacing: at most one request per second to a site, each feed read at its source's interval (15 minutes or more), no credentials sent. An editor can ask for a fresh read of the feeds, but a site is never asked again within a few minutes of its last read, and a site that asked the crawler to slow down or refused it is left alone until its next scheduled read.
- robots.txt: honoured for every feed and page (RFC 9309) and cached for an hour. An unreachable robots.txt, or one answering 401, 403 or 429, disallows the whole site.
- When a site refuses access (401, 403, 429 or a bot-challenge page): the crawler records the refusal, honours
Retry-After, backs off (5 minutes, doubling up to 6 hours), and after 5 refusals in a row stops requesting that site's articles, checking again once a day. It never tries to get around a refusal: no browser impersonation, no rotating addresses, no headless browser. The article then reaches editors only as the site's own feed describes it (headline, description and image), labelled as an excerpt with a link to the original.
Images
When a story uses an article's own image (the share image the page or feed names), ARCEL downloads that image once, with the same identity, to check that it suits the story; readers then see it linked from the publisher's site with its credit, never a copy. The download honours the image site's robots.txt exactly as above (a refusing or unreachable robots.txt means the image is not used), names the article page as its Referer, follows at most three redirects and accepts only JPEG, PNG, WebP and GIF images.
Allowlisting or excluding the crawler
Publishers who want the crawler allowlisted (for full articles under a recorded permission), or excluded entirely, can say so through their publisher contact at ARCEL. An excluded site is removed from the reviewed source list, and its stored content can be withdrawn.