SkimmrBot

About Skimmr

Skimmr provides content feeds built from crawled web content. Our customers are primarily members of the press.

As citizens of a democracy, we have a direct interest in journalists having broad, lawful access to the information they need — particularly information that public bodies are legally required to publish.

How we crawl

We crawl lawfully. We do not bypass logins, we do not access paywalled content, and we do not use credentials we are not entitled to use.

Public bodies and machine-readable access

Public bodies in Germany are subject to statutory publication and transparency obligations, and — under the Datennutzungsgesetz (DNG) and the Hamburgisches Transparenzgesetz, among others — to requirements that public information be made available in machine-readable form.

Where a public body publishes information it is legally obliged to publish, but configures its site to exclude automated access to that information, we consider that configuration to be in tension with those obligations.

In those cases we may contact the operator, set out our position, and ask that access for legitimate crawlers be permitted. In other cases or where that does not resolve the matter, we pursue it through the available legal channels, including requests under the applicable Informationsfreiheits- and Transparenzgesetze and complaints to the responsible supervisory authority.

We do not treat commercial or private websites this way. Their robots.txt is decisive.

About SkimmrBot

SkimmrBot is the public web crawler operated by Skimmr (Kordiam AI Systems GmbH, Hamburg, Germany). It visits publicly accessible web pages to build content feeds for Skimmr customers.

This page explains how to identify SkimmrBot, how to control or block it, how to verify a request claiming to be SkimmrBot, and how to request that your site is no longer crawled.

If you have a question this page doesn’t answer, write to info@skimmr.ai.

How to identify SkimmrBot

SkimmrBot identifies itself in the User-Agent header of every request:

SkimmrBot/1.0 (+https://skimmr.ai/bot)

The robots.txt user-agent token is:

SkimmrBot

How to control or block SkimmrBot

Place a robots.txt file at the root of your domain (e.g. https://example.com/robots.txt).

Block SkimmrBot from your entire site:

User-agent: SkimmrBot Disallow: /

Block SkimmrBot from specific areas only:

User-agent: SkimmrBot Disallow: /members/ Disallow: /internal/ Allow: /

Changes are picked up on SkimmrBot’s next visit.

How to reduce load caused by SkimmrBot

We attempt to crawl narrowly, targeting only the areas of a site relevant to our customers, and we limit request rates.

You can reduce load further by providing a well-formed sitemap (/sitemap.xml) or a regularly updated RSS feed. It may take days or weeks for SkimmrBot to detect it, but once it does, SkimmrBot prefers those sources.

If SkimmrBot is causing problems on your site, contact us at info@skimmr.ai and we will try to adjust.

Contact

Kordiam AI Systems GmbH
Falkenried 74a, 20251 Hamburg, Germany
info@skimmr.ai

Kordiam AI Systems GmbH, Falkenried 74a, 20251 Hamburg, Germany

German Impressum:

Geschäftsführer: Matthias Kretschmer

Phone: +49 40 88 14 170 - 0

Email: info@skimmr.ai

Kordiam AI Systems GmbH - VAT ID: DE455576115

AG Hamburg HRB 192338

Inhaltlich Verantwortlicher: Matthias Kretschmer

Datenschutzbeauftragter: dpo@kordiam.io