What it is
WorldFeedMonitor is a news aggregator. It reads the RSS and Atom feeds that publishers make public, and serves them to our customers as a structured API.
If you are reading this, you probably found the address in a User-Agent header in your logs. This page exists so you can see exactly what that request was and decide what you want to do about it.
How to identify it
Every request we make sends this User-Agent:
Mozilla/5.0 (compatible; WorldFeedMonitor/1.0; +https://staging.worldfeedmonitor.com/crawler)
We do not disguise our requests as a browser, and we do not rotate identities. A request that does not carry that string is not from us.
What it fetches
Only feed URLs — the RSS or Atom document itself. We do not crawl your site, follow links out of the feed, fetch article pages, or request images and assets. One feed URL is one request.
We store what the feed publishes: headline, summary, link, publication date, categories and, where the feed supplies one, the image URL. Our customers are always given the link back to your page.
How often
Each feed is read at most once every fifteen minutes. If a feed has not changed since the last read, we notice and do nothing further with it.
We do not fetch in parallel bursts against a single host. If a request is answered — including answered with a refusal — we take that as the answer and do not ask again until the next scheduled read, which a refusal makes a much longer wait. If a request is not answered at all, we may wait longer for the same server, from the same address, under the same name; that is the only kind of second attempt we make.
If you refuse it
A 403 or 429 is treated as your decision, not as an error to retry around. We slow down on it: after a few refusals we drop to hourly, then to every six hours, then to once a day. Nothing in our crawler responds to a refusal by changing our User-Agent, our address, or anything else to get a different answer.
We keep the feed on that slow schedule while there is still somewhere to ask from, because blocks are sometimes temporary or accidental. When every address we have has been refused, we stop reading that feed altogether and do not schedule it again. A block that holds from everywhere is an answer, and asking a nine hundredth time is not respecting it.
If your refusal is deliberate and permanent and you would rather not wait for us to work that out, tell us and we will remove the feed immediately.
One exception, stated plainly: a small number of feeds — about a hundred out of some two and a half thousand — are read through a second egress address, because a filter in front of them was refusing our server's address rather than our crawler. That is set per feed by a person, never automatically, and the User-Agent is identical either way. If yours is one of them and you would rather it were not, email us and we will stop.
That remains the only reason a feed is ever read from a second address, and it is still never a response to being refused: nothing switches address because it was turned away. What being refused from both does is end the matter — a feed refused on every address it has been read from is switched off.
How to block it
Any of these works, and none of them needs our cooperation:
- Disallow the feed path for
WorldFeedMonitorin yourrobots.txt. - Return
403to theUser-Agentabove. We will back off automatically. - Email us and we will stop, whatever your server says.
Allowlisting and contact
If you would rather we kept reading your feed and something in front of it is refusing us — a WAF rule or an address-reputation filter, usually — allowlisting the User-Agent above is enough.
Either way, write to legal@worldfeedmonitor.com. A request to stop is honoured without argument, and we do not ask for a reason.