If you are scraping Wildberries, Ozon, or any other site by loading full HTML pages, you are paying 5-10 times more for proxy traffic than you could. Each product page is 200-800 KB of markup, scripts, and styles, of which you only need a few fields: price, availability, rating. In this article, we will explore how to find the hidden API of a website and retrieve the same data directly in a compact JSON format.
Why HTML Parsing Consumes Proxy Traffic
When a scraper loads a page via a standard HTTP request or through a headless browser (Selenium, Puppeteer, Playwright), the server delivers a complete HTML document: markup, inline scripts, styles, sometimes base64 images, and hundreds of lines of JSON with data for advertising widgets that you do not need. The average product card on Wildberries weighs 300-600 KB, while on Ozon it can be up to 800 KB when counting all related resources (CSS, fonts, trackers).
If you monitor 10,000 products once a day through 3 proxy sessions, this easily adds up to tens of gigabytes of traffic per month. Residential and mobile proxies are usually sold by traffic, so every extra megabyte is a direct expense. Meanwhile, the actual data you need — price, discount, stock availability, rating — occupies 1-5 KB in the JSON response. This results in a difference of 100-200 times in volume for a single product, and considering the overhead costs of browser rendering, the savings in time and CPU are even greater.
An additional problem with HTML parsing is its fragility. Marketplace websites regularly change their layout, CSS classes, and DOM structure. Each such change breaks a parser built on XPath or CSS selectors. The internal API changes much less frequently because it affects the operation of both the mobile application and the website frontend simultaneously.
What is a Hidden API and Where Does It Come From
Almost every modern website is a SPA (Single Page Application) or hybrid application, where the browser first loads the "skeleton" of the page and then makes additional requests to the internal API for real data: prices, stock levels, reviews, recommendations via JavaScript. These requests are called hidden or internal APIs — they are not publicly documented but are fully visible in the browser's traffic.
Technically, these are usually REST or GraphQL endpoints that return data in JSON format. For example, on Wildberries, a product card is loaded through requests like card.wb.ru and wbx-content-v2.wbstatic.net, while prices and stock levels are fetched through a separate request to basket-01.wb.ru and similar domains. Ozon follows a similar logic: the frontend accesses the internal composer API, which aggregates data from microservices.
It is important to understand: using such an API is not technically hacking — you are simply repeating the same requests that a regular user’s browser makes. However, websites protect these endpoints through anti-bot systems, so careful imitation of real client behavior is necessary, including using quality proxies.
How to Find a Hidden API Using DevTools
You can find the internal API without writing a single line of code by using the built-in tools in Chrome or Firefox. Here’s a step-by-step algorithm:
- Open the desired product page in Chrome, press F12, and go to the Network tab.
- In the request filter, select the type Fetch/XHR — this will filter out image, font, and static resource loading.
- Refresh the page (F5) and look at the list of requests that appeared after loading the skeleton of the page.
- Find a request whose response (Response tab) shows the product price, name, or other necessary fields in JSON format.
- Click on this request and copy it as cURL (right-click → Copy → Copy as cURL) — this will give you the full set of headers, cookies, and parameters.
- Check which parameters in the URL are mandatory (product ID, region, API version) and which can be removed without losing data.
After that, it is enough to repeat this request using a standard HTTP library, substituting the necessary product ID instead of rendering the entire page. This works for most marketplaces — Wildberries, Ozon, Avito, as well as for many foreign platforms like Amazon and eBay.
Traffic Comparison: HTML vs JSON API
The difference in data volume is so significant that it is worth showing it in numbers. Below are average measurements for a single product card on popular marketplaces.
| Parsing Method | Average Response Size | Loading Time | JS Rendering Required |
|---|---|---|---|
| Full HTML via Selenium | 400-800 KB | 1.5-4 sec | Yes |
| Simple HTTP Request (requests) | 150-300 KB | 0.3-0.8 sec | No |
| Hidden JSON API | 3-15 KB | 0.1-0.3 sec | No |
When monitoring 50,000 products a day, switching from a headless browser to direct API requests reduces traffic from approximately 30-40 GB to 300-700 MB per month. This not only saves on proxy traffic but also reduces the load on the server infrastructure of the scraper — less CPU for rendering, less memory, and faster data collection.
Practical Example in Python
Let’s consider a simplified example: retrieving the price and stock level of a product via a direct request to the internal API instead of loading the full page. This is a template for educational purposes — the exact endpoints and parameters need to be determined through DevTools for a specific site, as the structure of requests may vary depending on the region and API version.
import requests
def get_product_data(product_id: str, proxies: dict = None) -> dict:
"""
Retrieves product data via the internal API instead of full HTML.
proxies — a dictionary with proxies in the format requests: {"http": "...", "https": "..."}
"""
url = f"https://card.example-marketplace.ru/v2/detail"
params = {
"nm": product_id,
"dest": "-1257786", # region, determined via DevTools
"spp": "0"
}
headers = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "
"AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/120.0 Safari/537.36",
"Accept": "application/json",
"Referer": f"https://www.example-marketplace.ru/catalog/{product_id}/detail.aspx"
}
response = requests.get(
url,
params=params,
headers=headers,
proxies=proxies,
timeout=10
)
response.raise_for_status()
data = response.json()
product = data["products"][0]
return {
"id": product["id"],
"name": product["name"],
"price": product["salePriceU"] / 100,
"stock": product.get("totalQuantity", 0),
"rating": product.get("reviewRating", None)
}
if __name__ == "__main__":
proxy = {
"http": "http://user:pass@proxy-host:port",
"https": "http://user:pass@proxy-host:port"
}
result = get_product_data("123456789", proxies=proxy)
print(result)
Note three points in this example. First, we specify the Referer header because many APIs check that the request comes "from a browser" and not directly via URL. Second, we use a realistic User-Agent, rather than the default from the requests library, which is easily detectable. Third, the entire request fits into a single HTTP call without rendering — this is what provides a multiple gain in traffic and speed.
For GraphQL endpoints, the logic is similar, but instead of GET parameters, you send a POST request with a request body in JSON format, explicitly listing the required fields — this further reduces the response size, as the server only returns the requested data.
Working with Proxies When Making API Requests
Even when switching to a compact JSON format, you still need proxies — marketplaces limit the number of requests from a single IP and ban accounts for abnormal activity. The right choice of proxy type directly affects the stability of the scraper.
For mass API scraping of marketplaces like Wildberries or Ozon, data center proxies are well-suited — they provide high speed and low traffic costs, which is critical for frequent requests to lightweight JSON endpoints. However, if a specific API is protected by stricter anti-bot measures and bans entire data center subnets, it is wiser to switch to residential proxies — they use real IP addresses of home users and are less likely to be blocked by subnet.
For APIs tied to mobile applications (some versions of endpoints on Avito or marketplaces only return data over mobile traffic), it may be necessary to connect through mobile proxies — they simulate traffic from real mobile operators and pass checks that block regular IPs.
When setting up proxies in the scraper, it is also important to stagger requests over time and use IP rotation — even a compact JSON request repeated 1000 times a minute from a single address will raise suspicion with the protection system. Set up a pool of several proxy sessions and distribute the load among them, adding random delays of 1-3 seconds between requests.
Pitfalls: Tokens, Signatures, Anti-bot
Hidden APIs are not always wide open. Some websites protect their endpoints with additional mechanisms that need to be considered when building a scraper.
- Temporary session tokens — some APIs require a preliminary request for a token, which is then passed in the header of subsequent requests and has a limited lifespan (usually 5-30 minutes).
- Request signature — request parameters are hashed on the client with a secret key from the page's JS code. This signature needs to be either reproduced manually by deciphering the algorithm or executed through a headless browser only during the token acquisition stage, and thereafter send lightweight requests directly.
- Rate limiting by IP and User-Agent — exceeding the request frequency temporarily blocks access to the site. This can be resolved by rotating proxies and implementing reasonable delays.
- Header fingerprinting — some systems check the complete set of headers (order, presence of Accept-Language, Sec-Fetch-*) and block requests with an "incomplete" set typical for scripts rather than browsers.
- Geo-dependence of data — prices and stock levels on marketplaces may vary by region, so it is important to pass the correct region/warehouse parameter in the request; otherwise, the data will be irrelevant.
If the API is closed with a request signature that is difficult to reproduce, a compromise option is to use a headless browser (Playwright, Puppeteer) only for intercepting network requests and extracting the ready JSON response, without parsing the DOM. This is slower than a direct HTTP request, but still faster and easier than full page layout parsing.
Checklist Before Launching the Scraper on a Hidden API
- The endpoint has been found via DevTools, copied as cURL, and tested in Postman or via requests.
- Mandatory request parameters (product ID, region, API version) have been identified and excess ones excluded.
- User-Agent, Referer, and Accept-Language headers have been set realistically.
- It has been checked whether a session token or request signature is required, and a method for obtaining them has been devised.
- Proxy rotation and random delays between requests have been configured.
- The appropriate type of proxy has been chosen for the specific website protection — data center, residential, or mobile.
- Error handling for 429 and 403 has been added with automatic switching to another proxy.
- Traffic volume logging has been set up to monitor actual savings.
Conclusion
Transitioning from parsing full HTML to working with hidden APIs is not just a technical optimization but a direct reduction in expenses for proxy traffic and infrastructure. Instead of loading hundreds of kilobytes of unnecessary markup, you receive compact JSON with exactly the fields needed for monitoring prices, stock levels, or ratings. An additional bonus is the scraper's resilience to changes in the website layout, as internal APIs change less frequently than the frontend.
However, the methodology for finding APIs does not eliminate the need for quality proxies — anti-bot systems of marketplaces closely monitor both HTML requests and calls to JSON endpoints. If you are monitoring Wildberries or Ozon in large volumes, start with fast data center proxies to reduce costs, and at the first signs of blocks, switch to residential or mobile IP pools for more stable scraper operation.