Does your Scrapy parser fail due to a timeout, the proxy pool depletes much faster than expected, and the logs are filled with 407 and 403 responses — does this sound familiar? In 90% of cases, the problem lies not with the proxies themselves, but with how the DownloaderMiddleware is written. We will discuss five of the most common mistakes in middleware that turn expensive traffic into junk requests and show how to fix them with code.
Mistake 1: Primitive rotation without considering proxy status
The most common construct found in tutorials is random.choice(PROXY_LIST) inside process_request. The problem is that such rotation does not know which proxy has just been banned and which is still alive. As a result, the parser continues to send requests through an already blocked IP, receives 403/429 responses, retries — and again selects the same address because the choice is random and does not exclude "bad" nodes.
The correct approach is to maintain the state of each proxy: the number of successful requests, the number of errors, and the time of last use. Here is a minimal working version:
import random
import time
class ProxyPool:
def __init__(self, proxies):
self.proxies = {p: {"fails": 0, "last_used": 0, "banned_until": 0} for p in proxies}
def get_proxy(self):
now = time.time()
available = [
p for p, state in self.proxies.items()
if state["banned_until"] < now
]
if not available:
# if all are banned, reset the oldest ban
available = list(self.proxies.keys())
return random.choice(available)
def mark_fail(self, proxy, cooldown=300):
self.proxies[proxy]["fails"] += 1
self.proxies[proxy]["banned_until"] = time.time() + cooldown
def mark_success(self, proxy):
self.proxies[proxy]["fails"] = 0
self.proxies[proxy]["last_used"] = time.time()
This pool excludes banned IPs during the "cooldown" period and returns them to circulation later. This already significantly reduces traffic consumption because you are not repeatedly hitting the same blocked node.
Mistake 2: Incorrect handling of retries and status codes
The second typical mistake is using the standard RetryMiddleware from Scrapy without modification. By default, it retries the request on the same proxy that failed it unless you explicitly mix in a proxy change in process_exception. As a result, we get the classic scenario: 3 retries, 3 bans, the request still fails, and the traffic has already been spent.
The second point is that not all status codes should be retried the same way. A 429 (Too Many Requests) requires a pause and an IP change, 403 usually indicates a ban on a specific proxy (immediate replacement is needed), while 5xx is often a temporary issue on the server side, and it can be retried on the same proxy. If everything is lumped together, the middleware either burns proxies too aggressively or waits too long where an IP change was needed immediately.
class SmartRetryMiddleware:
def __init__(self, pool):
self.pool = pool
def process_response(self, request, response, spider):
proxy = request.meta.get("proxy")
if response.status in (403, 407):
if proxy:
self.pool.mark_fail(proxy, cooldown=600)
new_request = request.copy()
new_request.meta["proxy"] = self.pool.get_proxy()
new_request.dont_filter = True
return new_request
if response.status == 429:
if proxy:
self.pool.mark_fail(proxy, cooldown=120)
new_request = request.copy()
new_request.meta["proxy"] = self.pool.get_proxy()
new_request.dont_filter = True
return new_request
if proxy:
self.pool.mark_success(proxy)
return response
Here it is important to differentiate cooldowns by error type: a hard ban (403/407) — long pause, rate limit (429) — short. This saves dozens of percent of traffic on long runs.
Mistake 3: Lack of sticky sessions for authenticated sites
If the parser works with a site that requires login, a shopping cart, pagination with state retention, or CAPTCHA with IP verification — changing proxies for each request breaks the session. The site sees that request #1 came from one IP, while request #2 (within the same cookie session) came from another, triggering bot protection instantly, even if both IPs are "clean."
The solution is to bind one proxy to one logical session (for example, to a specific account or to a chain of requests within the same domain) for a fixed time, rather than changing it for each request. This is called a sticky session.
class StickySessionMiddleware:
def __init__(self, pool, ttl=600):
self.pool = pool
self.ttl = ttl
self.sessions = {} # session_id -> (proxy, expires_at)
def process_request(self, request, spider):
session_id = request.meta.get("session_id")
if not session_id:
return
now = time.time()
session = self.sessions.get(session_id)
if session and session[1] > now:
request.meta["proxy"] = session[0]
else:
proxy = self.pool.get_proxy()
self.sessions[session_id] = (proxy, now + self.ttl)
request.meta["proxy"] = proxy
For tasks that require stable IP throughout the session — authentication, working with personal accounts, multi-step forms — the best fit are residential proxies with session support: they allow you to maintain the same outgoing IP for several minutes or hours and then change it in a controlled manner, rather than randomly for each request.
Mistake 4: Incorrect proxy authorization handling
The fourth mistake is technical but occurs in almost every second project. Developers pass the proxy username and password directly in the URL format http://user:pass@ip:port via request.meta["proxy"]. This works in most cases, but when using proxies through an HTTPS tunnel or when working with certain providers, this method of authorization is incorrectly handled by the standard HttpProxyMiddleware, and requests fail with 407 Proxy Authentication Required, even though the credentials are correct.
A more reliable way is to explicitly pass the Proxy-Authorization header, encoded in base64:
import base64
class ProxyAuthMiddleware:
def process_request(self, request, spider):
proxy = request.meta.get("proxy")
if not proxy:
return
# proxy without credentials in the URL
request.meta["proxy"] = proxy
user = request.meta.get("proxy_user")
password = request.meta.get("proxy_pass")
if user and password:
credentials = f"{user}:{password}"
encoded = base64.b64encode(credentials.encode()).decode()
request.headers["Proxy-Authorization"] = f"Basic {encoded}"
This approach works more reliably when scaling to hundreds of parallel requests and does not depend on how a specific version of the library parses URLs with embedded credentials. This is especially critical when working with mobile proxies, where authentication is often tied to an IP whitelist or strict header verification.
Mistake 5: No monitoring and logging of bans
The last and possibly most costly mistake is the lack of logging for which proxies are getting banned, how often, and on which domains. Without this data, it is impossible to understand what exactly is consuming traffic: whether the proxy pool has exhausted itself on a specific site or if the problem lies within the parser itself (too frequent requests, lack of delays, suspicious headers).
The minimum set of metrics that should be logged in the middleware includes:
- Number of requests per proxy during the parsing session
- Number and codes of errors (403, 407, 429, 5xx) for each proxy
- Lifetime of the proxy until the first ban
- Domains where bans occur most frequently
import logging
import json
logger = logging.getLogger("proxy_stats")
class ProxyStatsMiddleware:
def __init__(self):
self.stats = {}
def process_response(self, request, response, spider):
proxy = request.meta.get("proxy", "unknown")
domain = request.url.split("/")[2]
key = f"{proxy}|{domain}"
entry = self.stats.setdefault(key, {"requests": 0, "errors": 0})
entry["requests"] += 1
if response.status in (403, 407, 429):
entry["errors"] += 1
if entry["requests"] % 50 == 0:
logger.info(json.dumps(self.stats))
return response
Without such statistics, any attempts to "optimize" the middleware come down to guesswork. With it, you can clearly see: if, for example, a specific domain bans 80% of the proxy pool within the first 10 requests — the problem is not with the proxies, but with the request patterns (no User-Agent rotation, too high frequency, lack of delays between requests).
Working example of middleware in full
We gather everything together in settings.py — the order of middleware is critical because it determines the order in which checks are applied:
DOWNLOADER_MIDDLEWARES = {
"myproject.middlewares.ProxyAuthMiddleware": 350,
"myproject.middlewares.StickySessionMiddleware": 400,
"myproject.middlewares.SmartRetryMiddleware": 550,
"myproject.middlewares.ProxyStatsMiddleware": 900,
}
RETRY_ENABLED = False # disable standard retry, use your own
DOWNLOAD_TIMEOUT = 15
CONCURRENT_REQUESTS_PER_DOMAIN = 8
Note the RETRY_ENABLED = False — this is crucial; otherwise, Scrapy's built-in retry mechanism will conflict with your proxy-switching logic, leading to duplicate or double-retried requests. It is also important to limit CONCURRENT_REQUESTS_PER_DOMAIN — too high concurrency on one domain, even with proxy rotation, looks suspicious to anti-bot systems.
Which type of proxy to choose for Scrapy
Middleware solves half the problem; the other half is the correct choice of the proxy pool for the task. Below is a comparison for typical parsing scenarios.
| Type of Proxy | When to Use | Pros | Cons |
|---|---|---|---|
| Datacenter Proxies | Mass scraping of open pages without strict anti-bot protection | High speed, low traffic cost | Easily detected, often blacklisted |
| Residential Proxies | Scraping marketplaces, sites with JS protection, authentication | Real IPs, low ban rates, support for sticky sessions | Slower speed compared to DC |
| Mobile Proxies | Scraping mobile versions of sites and APIs with strict anti-bot protection | Maximum trust from target sites | Highest traffic cost |
A practical rule: for scraping static pages without strong protection, cheap datacenter proxies combined with well-designed middleware from this article are suitable. If the site uses Cloudflare, PerimeterX, DataDome, or similar systems — datacenter IPs will be banned almost immediately, and it is more advantageous to switch to a residential pool right away, saving time on debugging middleware.
Checklist before launching the parser in production
- Middleware excludes banned proxies during cooldown, rather than randomly selecting them
- Retry logic differentiates error types (403/407 vs 429 vs 5xx) with different cooldowns
- For session tasks, sticky proxy binding is used instead of changing it for each request
- Proxy authorization is passed through the Proxy-Authorization header, not just through the URL
- Statistics on bans by domains and proxies are maintained for diagnosing problems
- Standard Scrapy RetryMiddleware is disabled to avoid conflicts with custom logic
- Concurrency is limited to reasonable values, not set to maximum
Conclusion
Most issues with proxy traffic consumption in Scrapy can be resolved not by purchasing a larger pool of IPs, but by fixing the middleware logic: differentiating errors by type, using sticky sessions for complex scenarios, proper authorization, and constant monitoring of bans. The code from this article can be taken as a basis and adapted for a specific project — the structure remains functional for both small parsers and distributed Scrapy Cluster installations.
If the middleware is already set up correctly, and bans still occur too frequently — the issue is likely with the quality of the IP pool itself. For scraping sites with advanced anti-bot protection, it is worth trying residential proxies with session support — they significantly reduce the rate of protection triggers compared to datacenter addresses while maintaining the same middleware logic.