SkimmrBot
About Skimmr
Skimmr provides content feeds built from crawled web content. Our customers are primarily members of the press.
As citizens of a democracy, we have a direct interest in journalists having broad, lawful access to the information they need — particularly information that public bodies are legally required to publish.
How we crawl
We crawl lawfully. We do not bypass logins, we do not access paywalled content, and we do not use credentials we are not entitled to use.
Public bodies and machine-readable access
Public bodies in Germany are subject to statutory publication and transparency obligations, and — under the Datennutzungsgesetz (DNG) and the Hamburgisches Transparenzgesetz, among others — to requirements that public information be made available in machine-readable form.
Where a public body publishes information it is legally obliged to publish, but configures its site to exclude automated access to that information, we consider that configuration to be in tension with those obligations.
In those cases we may contact the operator, set out our position, and ask that access for legitimate crawlers be permitted. In other cases or where that does not resolve the matter, we pursue it through the available legal channels, including requests under the applicable Informationsfreiheits- and Transparenzgesetze and complaints to the responsible supervisory authority.
We do not treat commercial or private websites this way. Their robots.txt is decisive.
About SkimmrBot
SkimmrBot is the public web crawler operated by Skimmr (Kordiam AI Systems GmbH, Hamburg, Germany). It visits publicly accessible web pages to build content feeds for Skimmr customers.
This page explains how to identify SkimmrBot, how to control or block it, how to verify a request claiming to be SkimmrBot, and how to request that your site is no longer crawled.
If you have a question this page doesn’t answer, write to info@skimmr.ai.
How to identify SkimmrBot
SkimmrBot identifies itself in the User-Agent header of every request:
SkimmrBot/1.0 (+https://skimmr.ai/bot)
The robots.txt user-agent token is:
SkimmrBot
How to control or block SkimmrBot
Place a robots.txt file at the root of your domain (e.g. https://example.com/robots.txt).
Block SkimmrBot from your entire site:
User-agent: SkimmrBot Disallow: /
Block SkimmrBot from specific areas only:
User-agent: SkimmrBot Disallow: /members/ Disallow: /internal/ Allow: /
Changes are picked up on SkimmrBot’s next visit.
How to reduce load caused by SkimmrBot
We attempt to crawl narrowly, targeting only the areas of a site relevant to our customers, and we limit request rates.
You can reduce load further by providing a well-formed sitemap (/sitemap.xml) or a regularly updated RSS feed. It may take days or weeks for SkimmrBot to detect it, but once it does, SkimmrBot prefers those sources.
If SkimmrBot is causing problems on your site, contact us at info@skimmr.ai and we will try to adjust.
Contact
Kordiam AI Systems GmbH
Falkenried 74a, 20251 Hamburg, Germany
info@skimmr.ai
Kordiam AI Systems GmbH, Falkenried 74a, 20251 Hamburg, Germany
German Impressum:
Geschäftsführer: Matthias Kretschmer
Phone: +49 40 88 14 170 - 0
Email: info@skimmr.ai
Kordiam AI Systems GmbH - VAT ID: DE455576115
AG Hamburg HRB 192338
Inhaltlich Verantwortlicher: Matthias Kretschmer
Datenschutzbeauftragter: dpo@kordiam.io