A search crawler comes to a site with a limited budget of requests. Every request that hits a nonexistent address is wasted: the page it came for stays unread
Over two days our logs recorded 129 errors in total. By crawler they split like this:
| crawler | errors |
|---|---|
| yandexbot | 82 |
| bingbot | 23 |
| amazonbot | 13 |
More than half come from one crawler. We look at what exactly it requests, and the picture becomes unambiguous: the most frequently requested address looks like "/kz/almaty/study-abroad/globaleducation.kz/ky/"
The "/ky/" tail is a Kyrgyz language version. We do not build one. Our language versions live at "/ru/", "/en/" and "/qz/", and each of them responds with code 200
So the crawler knocks on an address we never had. It took it from the markup of the language versions, where the language code turned out to be written in another standard
82 requests in two days are 82 business pages that the crawler could have read and did not. With a thousand-plus catalog pages, such a leak eats a noticeable share of the daily budget
One redirect line on the web server side: an address with the "/ky/" tail answers with a redirect to "/qz/". The crawler gets a ready page instead of an error, the budget stops leaking, and the markup is fixed separately and without hurry
Only the server log shows a crawler its errors. They do not appear in position reports, and in the webmaster panel they surface with a delay. There is one useful habit: once a week, look at the top addresses in the error log and ask where the crawler got that address