ooon.ai
all articles

The crawler knocks on an empty door: 129 errors in two days and one line that removes them

A search crawler comes to a site with a limited budget of requests. Every request that hits a nonexistent address is wasted: the page it came for stays unread

What did the measurement show?

Over two days our logs recorded 129 errors in total. By crawler they split like this:

crawlererrors
yandexbot82
bingbot23
amazonbot13

More than half come from one crawler. We look at what exactly it requests, and the picture becomes unambiguous: the most frequently requested address looks like "/kz/almaty/study-abroad/globaleducation.kz/ky/"

Where did the address come from?

The "/ky/" tail is a Kyrgyz language version. We do not build one. Our language versions live at "/ru/", "/en/" and "/qz/", and each of them responds with code 200

So the crawler knocks on an address we never had. It took it from the markup of the language versions, where the language code turned out to be written in another standard

The price of the problem

82 requests in two days are 82 business pages that the crawler could have read and did not. With a thousand-plus catalog pages, such a leak eats a noticeable share of the daily budget

What fixes this problem?

One redirect line on the web server side: an address with the "/ky/" tail answers with a redirect to "/qz/". The crawler gets a ready page instead of an error, the budget stops leaking, and the markup is fixed separately and without hurry

What does this have in common for any site?

Only the server log shows a crawler its errors. They do not appear in position reports, and in the webmaster panel they surface with a delay. There is one useful habit: once a week, look at the top addresses in the error log and ask where the crawler got that address