← Back to Blog

Parser Retry Logic: How Retries Consume Up to 40% of Traffic and How to Fix It

Repeated requests can consume up to 40% of proxy traffic. We show how to calculate actual losses and reduce them through proper backoff, circuit breaker, and smart proxy rotation.

📅September 27, 2026

If the proxy bill is growing faster than the amount of data collected, the problem is almost always in the retry logic. The parser silently repeats failed requests several times, wasting traffic on timeouts and CAPTCHAs, while the developer doesn't even see these expenses in the logs. Let's analyze how to calculate real losses and reduce them without compromising data quality.

Why retry logic consumes traffic

Most parsers are written with naive retry logic: if a request fails — repeat it, and so on for 3-5 times. The problem is that each retry request is not only a new HTTP request but also a complete cycle: TCP handshake, TLS negotiation, loading the entire page (even if only one block of data is needed), and sometimes even reloading images or JS files if the parser uses a headless browser instead of a simple HTTP client.

Retries are particularly costly when working through residential proxies, where traffic is charged by volume rather than by the number of requests. One failed request to a product page on Wildberries with images and scripts can cost 300-500 KB. If the parser makes 3 retries on a timeout, you pay for the same failed request four times in a row — and this is without any guarantee that the fourth attempt will be successful.

The second reason is retries on errors that cannot be fixed by repeating. If the site returns a 403 due to bot detection, a repeated request with the same fingerprint and the same session cookies will almost certainly receive the same response. The parser wastes traffic on attempts that mathematically cannot succeed until the IP address or browser fingerprint changes.

How much traffic is actually spent on retries

To understand the scale of the problem, let's take a simple example. The parser collects product cards from a marketplace, with an average response size of 250 KB (HTML + JSON API + partially static content). With stable operation without blocks, the share of failed requests remains at 5-8%. However, with aggressive parsing through cheap data center proxies, this figure can rise to 25-35% because the target provider quickly recognizes the pattern and starts returning CAPTCHAs or temporary IP bans.

Let's calculate with numbers. Suppose we need to collect 100,000 product cards:

Share of failed requests Retries per request (on average) Total traffic Overrun
5% 0.15 28.75 GB +15%
15% 0.45 36.25 GB +45%
30% 0.90 47.5 GB +90%

As can be seen, with a failure rate of 30% and a retry strategy of "retry up to 3 times," the actual traffic almost doubles compared to the theoretical minimum of 25 GB. These extra 20+ GB are direct losses to the proxy budget, which can be reduced by rethinking the retry logic.

Common mistakes in parser retry logic

Before fixing the retry logic, it is worth recognizing the typical anti-patterns that occur in 90% of custom parsers:

  • Retry without analyzing the error code. A retry is triggered for any abnormal situation — 403, 429, 500, timeout, connection drop — although the handling strategy for them should differ.
  • Fixed delay between retries. For example, 2 seconds between attempts regardless of whether it is the first attempt or the fifth — this is either too aggressive for the site or too slow for large volumes.
  • Retry with the same IP and the same session. If the site has blocked the request due to fingerprinting, retrying with identical parameters does not change the outcome but wastes traffic.
  • No upper limit on attempts. Some parsers get stuck on "dead" URLs and make dozens of retries before giving up.
  • No distinction between temporary and permanent errors. 404 (page does not exist) and 503 (server temporarily unavailable) require different logic — but are often handled the same way.

Exponential backoff with Python code

A simple but effective solution is exponential delay with jitter (random variation), which reduces the number of pointless retries and spreads the load over time. Instead of a fixed pause between attempts, the delay grows exponentially, giving the site time to "cool down" after a block, and the parser itself does not waste traffic on requests that are almost guaranteed to fail.

import time
import random
import requests

def fetch_with_backoff(url, proxies, max_retries=4, base_delay=1.0):
    retryable_codes = {429, 500, 502, 503, 504}
    non_retryable_codes = {404, 410}

    for attempt in range(max_retries + 1):
        try:
            response = requests.get(url, proxies=proxies, timeout=10)

            if response.status_code == 200:
                return response

            if response.status_code in non_retryable_codes:
                # no point in retrying — the page physically does not exist
                return None

            if response.status_code not in retryable_codes:
                return None

        except (requests.exceptions.Timeout,
                requests.exceptions.ConnectionError):
            pass  # temporary network error — can retry

        if attempt == max_retries:
            return None

        # exponential delay with jitter
        delay = base_delay * (2 ** attempt) + random.uniform(0, 1)
        time.sleep(delay)

    return None

The key idea of this code is to categorize errors into three groups: those that cannot be fixed by retrying (404, 410), those that can be retried with a delay (429, 500-504, timeouts), and everything else, which is immediately considered a failure without incurring additional attempts. This categorization alone reduces unnecessary traffic by 20-30% compared to naive "retry everything."

Tip: add the Retry-After header to your handling — many sites provide hints on how many seconds to wait before retrying. Ignoring this header is a common cause of unnecessary bans and traffic.

Circuit breaker: when to stop

Exponential backoff helps at the level of a single request, but does not protect against situations where an entire domain or a specific proxy node is temporarily unavailable for hundreds of URLs in a row. Here, the circuit breaker pattern is needed — an "automatic switch" that monitors the error rate over the last period and temporarily halts attempts if it exceeds a threshold, instead of continuing to pound on a closed door.

class CircuitBreaker:
    def __init__(self, failure_threshold=0.5, window_size=50, cooldown=60):
        self.failure_threshold = failure_threshold
        self.window_size = window_size
        self.cooldown = cooldown
        self.results = []
        self.open_until = 0

    def is_open(self):
        return time.time() < self.open_until

    def record(self, success: bool):
        self.results.append(success)
        if len(self.results) > self.window_size:
            self.results.pop(0)

        if len(self.results) == self.window_size:
            failure_rate = 1 - sum(self.results) / self.window_size
            if failure_rate > self.failure_threshold:
                self.open_until = time.time() + self.cooldown
                self.results.clear()

The logic is simple: if more than half of the last 50 requests have failed, the parser suspends attempts for that domain or proxy for 60 seconds. During this time, you can change the IP, reduce the request speed, or switch to another proxy pool. This is especially important when working with target sites that temporarily ban the IP range after exceeding the request frequency — continuing to bang on a closed door means simply wasting traffic.

Smart proxy rotation during retries

One of the most effective measures against unnecessary retries is not to repeat a request from the same IP that has already received a rejection. The logic is simple: if the error is related to an IP block (403, 429, redirect to CAPTCHA), changing the proxy before retrying dramatically increases the chance of success and reduces the number of attempts.

For parsing marketplaces like Wildberries, Ozon, or Avito, a combination works well: regular requests go through data center proxies — they are faster and cheaper, and as soon as the block detection triggers several times in a row, the parser switches to residential proxies, which are less likely to be caught by anti-bot systems. This hybrid approach reduces overall traffic consumption because expensive residential IPs are used only where truly necessary, rather than for all requests in a row.

Error type Retry strategy Is IP change needed?
Connection timeout Backoff, 1-2 retries No
403 / CAPTCHA Immediate rotation Yes, mandatory
429 (rate limit) Backoff based on Retry-After Desirable
500-503 Backoff, 2-3 retries No
404 / 410 No retries —

For parsers that emulate mobile traffic (for example, collecting data from mobile versions of marketplace applications or social networks), it makes sense to use mobile proxies — they are less likely to raise suspicion with protection systems precisely because telecom operators assign the same IPs to thousands of real users simultaneously, making targeted blocking impractical for the target site.

Monitoring retry metrics

Without metrics, optimizing retry logic turns into guesswork. The minimum set of indicators to log with each request:

  • Retry rate — the share of requests that required at least one retry.
  • Success after retry — what percentage of retries ultimately succeeded (if this figure is low, retries are just burning traffic).
  • Traffic per successful result — the total volume of data transferred divided by the number of successfully collected records. This is a key metric of efficiency.
  • Error distribution by codes — helps to understand where the main traffic leak occurs: timeouts, 403, 429, or something else.
  • Retry rate by specific proxy nodes — if one IP has a retry rate of 80%, while others are 10%, the problem is local and can be solved by changing that specific node, not the entire logic.

Even a simple table in Google Sheets or a log in CSV with these five metrics, updated hourly, provides enough data to spot anomalies and timely adjust the strategy — for example, to reduce the request frequency to a specific section of the site or increase the share of residential IPs in the pool.

Traffic optimization checklist for retries

  1. Separate error codes into retryable and non-retryable — do not retry 404/410.
  2. Implement exponential backoff with jitter instead of fixed delays.
  3. Respect the Retry-After header if the site sends it.
  4. Change IP before retrying on 403 and suspected bot detection.
  5. Set a hard limit on the number of attempts (usually 3-4 is sufficient).
  6. Implement a circuit breaker for domains and proxy nodes with a high failure rate.
  7. Log retry rate and traffic per successful result — without metrics, optimization is impossible.
  8. Separate the proxy pool: cheap data centers for stable areas, residential or mobile IPs for problematic ones.

Conclusion

Retry logic is not a minor detail of the parser, but one of the main factors affecting the cost of data collection. A naive strategy of "retry everything" can increase actual traffic by 40-90% compared to the theoretical minimum, while most retries end in the same failure as the first attempt. Categorizing errors by type, exponential backoff, circuit breakers, and smart IP rotation allow for significant reductions in these losses without compromising the completeness of the collected data.

If your parser works with sites that aggressively detect bots — marketplaces, social networks, advertising platforms — it is worth combining several types of proxies depending on the task. For basic operations, data center proxies will suffice, while for maximum resistance to blocks, residential proxies with real IP addresses of ordinary users are needed. This hybrid approach, combined with smart retry logic, results in a noticeable reduction in traffic consumption while maintaining the same volume of collected data.