News generation on subscription · newsrooms, blogs and companies

Crawler

NaitoBot

NaitoBot is the user agent Naito uses to read a news article when it writes an original piece based on it. This page explains what it does, what it does not do, and how to control or block it.

How to recognize it

Requests identify themselves with this user agent:

NaitoBot/0.1 (+https://naito.news/bot)

For the small share of pages that only render with JavaScript, a headless Chromium fallback is used after the same robots.txt check. Those requests currently present a standard browser user agent instead of NaitoBot.

What it fetches

Only individual article URLs: the ones reported by the news data APIs Naito subscribes to, or the URLs a Naito customer submits for an on-demand article. NaitoBot does not crawl sites: it does not follow links, read sitemaps or feeds, or build an index of your pages.

It never logs in to a site and never uses credentials.

What happens to the page

Naito extracts the facts from the article and writes a new, original text for its customer, in the customer’s own voice. It is not designed to republish or closely paraphrase your text, and the source domain is credited. To let a reviewer verify each claim, a customer’s review queue (and, with the WordPress plugin, a collapsed “Sources” section) can show short passages from the source.

How it behaves

  • It reads your robots.txt before fetching and obeys rules for the NaitoBot token.
  • If robots.txt cannot be retrieved, it treats the site as allowed.
  • At most 2 simultaneous requests per domain, with an 8-second timeout per request.

How to block it

Add this to your robots.txt. It applies from our next fetch, since robots.txt is re-read on every run:

User-agent: NaitoBot
Disallow: /

Blocking NaitoBot stops Naito from fetching your pages. Headlines and short descriptions that third-party news data APIs list about your site are outside our control.

Contact

Questions, or something that looks wrong in your logs? Use the support form and mention NaitoBot.