You have purchased an expensive residential IP, set up rotation, and inserted a realistic User-Agent — yet the parser still gets blocked by a CAPTCHA or receives an empty response. The problem is almost always not with the IP, but with the TLS fingerprint: the library you are using to send the HTTPS request "sounds" different from a real browser. Anti-bot systems of Wildberries, Ozon, Cloudflare, and Akamai detect this before they even check your IP address.
What is TLS fingerprint and why it is more important than IP
When a client establishes an HTTPS connection, it sends a ClientHello packet to the server — part of the TLS handshake. It contains an encrypted list of supported TLS versions, cipher suites, extension order, elliptic curves, and compression algorithms. This set of parameters is unique for each combination of "library + operating system + TLS stack version."
Chrome, Firefox, and Safari generate ClientHello in their own way, and this set hardly changes from request to request — unlike IP or User-Agent, which can be easily spoofed in text. However, standard HTTP libraries — requests, urllib3, the standard HttpClient in Java, and the built-in TLS stack of Node.js — create a completely different ClientHello because they use OpenSSL or another library differently than a browser.
This is why you can connect a perfectly "clean" residential IP, insert a fresh User-Agent from a real Chrome — and still get blocked. The server sees the user's IP from a residential area, sees the header "Chrome 124," but the TLS handshake says: "this is a Python script." The mismatch is a direct signal for the anti-bot.
How anti-bot systems detect the parser by JA3/JA4
To convert the parameters of ClientHello into a compact identifier, the JA3 algorithm (and its newer version JA4) is used. It takes the TLS version, the list of ciphers, extensions, and curves, concatenates them into a string, and hashes them using MD5. This results in a short hash like 769,47-53-5-10...,0-23-65281...,29-23-24,0, which uniquely identifies the client's "fingerprint."
Anti-bot providers (Cloudflare, Akamai, PerimeterX, DataDome — and their analogs used by Wildberries and Ozon) maintain databases of known JA3/JA4 hashes of popular HTTP libraries: requests, aiohttp, Scrapy, Node fetch, Java HttpClient, Go net/http. If the hash matches a known "script" signature rather than a Chrome/Firefox/Safari signature — the request is flagged as suspicious even before behavior analysis.
Next, the system checks the match of the TLS fingerprint with the declared User-Agent. If the headers say "Chrome 124 on Windows," but the TLS fingerprint corresponds to OpenSSL 1.1.1 from the standard Python library — this is called TLS/HTTP mismatch, one of the most reliable signals of automation detection. This is how parsers are detected even with a perfect residential IP and correct headers.
How to check your TLS fingerprint: tools
Before fixing the problem, you need to see what the server sees. There are several public services that show your JA3/JA4 hash and the complete set of ClientHello parameters:
- tls.peet.ws — shows JA3, JA4, a list of ciphers and extensions in JSON format, convenient for automated script checking.
- ja3er.com — a database of known JA3 hashes linked to specific libraries and browsers.
- browserleaks.com/tls — a visual comparison of your fingerprint with typical browser fingerprints.
- Wireshark locally — if you want to see the raw ClientHello packet when sending a request from your script.
A practical test is simple: open tls.peet.ws in regular Chrome and note the JA4 hash. Then send a GET request to the same address from your parser (via requests, curl_cffi, or any other library) through the same proxy and compare the hashes. If they differ — the server sees the difference between the "browser" and the "script" with every request, regardless of how clean the IP is.
Checking with Python: requests, httpx, curl_cffi
Let's analyze in practice why standard Python libraries expose the parser. A regular request via requests:
import requests
resp = requests.get("https://tls.peet.ws/api/all", proxies={
"https": "http://user:pass@proxy_host:port"
})
print(resp.json()["tls"]["ja4"])
# The result will differ from the JA4 of real Chrome,
# as requests uses the standard ssl module of Python
The problem is that requests and httpx use the system OpenSSL through the ssl module, and the order and set of TLS extensions are hard-coded and do not match Chrome/Firefox. The solution is the curl_cffi library, which uses a patched curl with real browser TLS profiles:
from curl_cffi import requests as cffi_requests
resp = cffi_requests.get(
"https://tls.peet.ws/api/all",
impersonate="chrome124",
proxies={"https": "http://user:pass@proxy_host:port"}
)
print(resp.json()["tls"]["ja4"])
# The hash will be identical to real Chrome 124 on desktop
The impersonate parameter forces curl_cffi to reproduce not only the ClientHello but also the order of HTTP/2 headers (frame order), which is also part of the fingerprint. A similar approach is used by the tls-client library for Go and undetected-chromedriver for those who parse through a real browser rather than through an HTTP client.
If parsing is done through a headless browser (Playwright, Puppeteer, Selenium), the TLS fingerprint is formed by the Chromium/Firefox engine and by default matches the real browser. But here another problem arises — automation signatures at the JS level (webdriver flags, canvas fingerprint), so for headless scenarios, additional patches like playwright-stealth are needed.
TLS + HTTP/2 + headers: why the combination is important
The TLS fingerprint is just one layer of detection. Anti-bot systems check several levels simultaneously:
- TLS ClientHello (JA3/JA4) — the set of ciphers and extensions.
- HTTP/2 fingerprint — the order of pseudo-headers (:method, :path, :authority), SETTINGS frame settings, window size.
- HTTP headers — the order and set of regular headers (Accept-Language, Sec-Ch-Ua, Sec-Fetch-*).
- User-Agent — must match the version of the TLS profile: if UA says "Chrome 124," but TLS corresponds to Chrome 110, this is also suspicious.
A common mistake is to update the User-Agent to the latest version of Chrome without updating the TLS profile in curl_cffi or another library. Such version discrepancies are visible to the anti-bot as clearly as a complete lack of masking. Ensure that the impersonate version and the version in the User-Agent match, and update both parameters synchronously when new browser versions are released.
Another point is the order of headers. The browser sends headers in a strictly defined order, while many HTTP libraries sort them alphabetically or by the order they were added in the code. Even if the set of headers is identical to that of the browser, the incorrect order is an additional signal for advanced anti-bot systems like DataDome.
The role of proxies: why a clean IP does not save you
A residential IP solves a specific task — it reduces suspicion based on geography, ASN, and address reputation. Data center IPs are often blacklisted because they generate automated traffic en masse, while residential IPs belong to real providers and ordinary users. For scraping Wildberries, Ozon, or Avito, this is critical: without a clean IP, requests are blocked based solely on this criterion, without even checking TLS.
However, IP and TLS fingerprint are two independent layers of protection, and they solve different problems. The IP tells the server "where" the request came from, while the TLS fingerprint tells "how" it was sent. Therefore, a combination of a clean IP and the correct TLS profile is the minimum requirement for stable scraping. For tasks with a high request frequency and aggressive anti-bot measures, it is better to use residential proxies, as they provide a low percentage of bans based on IP reputation, but be sure to combine them with a library that correctly reproduces the TLS profile of a real browser.
For monitoring prices on marketplaces, where speed and volume of requests are important, data center proxies are often used in combination with TLS masking through curl_cffi — this is cheaper than residential proxies and quite effective if the site's anti-bot system is not too aggressive. For tasks where the site actively checks mobile networks (for example, scraping mobile versions of applications via API), mobile proxies are used — they provide an additional level of trust due to the reputation of the operator networks.
Parser configuration checklist without detection
Assemble the checks into a single process before launching the parser in production:
- Measure the JA4 hash of your script through tls.peet.ws and compare it with a real browser of the same version.
- Use a library that supports TLS impersonation: curl_cffi (Python), tls-client (Go), CycleTLS (Node.js).
- Synchronize the TLS profile version (impersonate) with the version in the User-Agent.
- Check the order of HTTP headers — it should match that of a real browser, not be alphabetical.
- Connect a clean residential or mobile IP based on the geo of your task.
- Set up IP rotation separately from the TLS profile — do not rigidly tie one to the other.
- Regularly update the TLS profile when new versions of Chrome are released — old signatures get into anti-bot databases faster than it seems.
- For scenarios with JS checks (Cloudflare Challenge), use a headless browser with stealth patches instead of a clean HTTP client.
Comparison of libraries and tools
| Tool | Browser TLS fingerprint | Speed | When to use |
|---|---|---|---|
| requests / httpx | No, exposes script | High | Sites without TLS detection, internal APIs |
| curl_cffi | Yes, exact copy | High | Marketplaces, anti-bot Cloudflare/Akamai |
| tls-client (Go) | Yes | Very high | High load, mass scraping |
| Playwright / Puppeteer | Yes, real engine | Low | JS render, Cloudflare Challenge, complex SPAs |
| Scrapy (standard) | No | High | Sites without strict anti-bot protection |
Conclusion
The TLS fingerprint is a layer of protection that many parsers completely ignore, spending resources on finding the perfect IP and User-Agent, but forgetting that the very structure of the TLS handshake reveals automation before the server looks at the headers. The solution is to use libraries that support TLS impersonation (curl_cffi, tls-client), synchronize the profile version with the User-Agent, and check the final JA4 hash before launching at scale.
The IP remains an important factor — without a clean address, even a perfect TLS fingerprint will not help bypass network reputation blocks. For scraping marketplaces and monitoring prices, it makes sense to combine the correct TLS configuration with residential proxies — this combination covers both layers of detection and significantly reduces the percentage of bans during long scraping sessions.