The parser pulls 400 KB of HTML for eight fields that the site returns to itself in a JSON response of 8 KB. A difference of fifty times is not about "beautiful code"; it's about the bill for residential proxies, where you pay for every gigabyte. Let's explore how to find the internal API of a website, what prevents its replication in 2026, and when it's worth abandoning this idea.
Why look for a hidden API if HTML is already being parsed
Almost any modern interface—React, Vue, Angular, Next.js—first loads the page skeleton and then pulls data through separate requests to its own endpoints. These endpoints are undocumented, but they exist, respond with clean JSON, and are accessible without a headless browser.
What you gain by switching to them:
- Traffic drops significantly. In parsing a typical product listing, the HTML page weighs about 400 KB along with markup, styles, and trackers, while the corresponding JSON endpoint is about 8 KB, and it has more fields: internal IDs, stock levels, product variants.
- No browser needed. JavaScript rendering is eliminated, along with memory usage, CPU load, and dozens of additional requests for fonts and analytics.
- Data is already structured. No selectors that break due to changes in CSS classes.
- Fewer requests mean fewer chances for bans. Rendering a single catalog page in a browser involves dozens of requests to the site; the same amount of data through the API is just one.
For a project using residential proxies, this is a direct saving: the rate is calculated by gigabytes, and switching from rendering to JSON usually compresses the bill more than any tricks with blocking images. A related topic is how to reduce parser traffic by 5 times using other methods.
Step by step: how to find the endpoint
- First, check if there is an official API. Look at
/developers,/api,/docsof the target website. A public documented API is versioned and warns about deprecations—private APIs change silently. - Open DevTools (F12) and go to the Network tab, ensuring that recording is enabled.
- Enable the Fetch/XHR filter. This filters out images, fonts, and analytics, leaving only data requests.
- Clear the list to remove noise from the initial load.
- Trigger the desired data: scroll through the results, click "next page," apply a filter, or open a product card. The request you are interested in will appear at the moment of action.
- Find the response with your data. The fastest way is to use Ctrl+F in the Network panel: search for a unique value you see on the screen (SKU, exact price, part of the name) and see which request generated it.
- Copy the request in full: right-click on the line → Copy → Copy as cURL. Then convert it to code using curlconverter—this way, you won't lose any headers.
Typical paths to look at first include: /api/, /v1/, /v2/, /search, /products, /listings, /graphql.
Special case: sites on Next.js
Here, data often doesn't require a separate request at all—they are embedded directly in the HTML. On the old Pages Router, this is the block __NEXT_DATA__. On the App Router (Next.js 13 and newer), hydration data is spread across calls to self.__next_f.push() in several script nodes—this is the serialized payload of React Server Components. Parsing it manually is unpleasant: chunks reference each other through $ prefixes and can be cut off in the middle of a string. For Python, there is a library called nextflight that parses both the Flight payload from HTML and the raw RSC response (request with the header RSC: 1), suggesting to search for keys by name rather than by array indices—this way, the parser survives the redeployment of the site.
Reversing parameters: pagination and filters
The found endpoint is almost always parameterized. Three schemes are common:
- By pages:
?page=3&per_page=20 - Offset and limit:
?offset=40&limit=20 - Cursor:
?after=<token>&limit=20—the token for the next page comes in the body of the previous response
Three rules that save hours of debugging:
- Stop at an empty batch rather than a pre-calculated number of pages: the
totalcounter in private APIs lies more often than one would like. - Check the actual size of the batch. If you requested 100 and received 20, then the endpoint has its own ceiling, and your pagination arithmetic is already incorrect.
- Don't go to page 500. Deep pagination is often cut off by the server; instead, slice the selection with filters—by category, by price range, by date.
Why cURL from the browser works, but your code does not
This is the most common point of failure, and the reason is almost always the same: a missing header. The copied cURL carries the entire context of the request, while a custom client does not.
What usually turns out to be mandatory:
- Custom headers with the
X-prefix—X-CSRF-Token,X-Requested-With: XMLHttpRequest, and variousX-*-Tokenthat the frontend inserts itself. Without them, you will receive a response in the 400–500 range. Referer—a contextual header generated by user action. Many endpoints check that the request "came from its own page."Authorization: Bearer <JWT>—a short-lived token, usually lasting 15–60 minutes. Hardcoding it is pointless: you need to be able to obtain a fresh one.- Session cookies—keep them in a session object, rather than copying them manually.
- Correct
Content-Typefor POST:application/jsonandapplication/x-www-form-urlencodedencode the body differently, and a mismatch with the declared type silently breaks the request.
Where to find the tokens themselves if they are not in cookies: in the HTML source inside <script> (search for a known value using Ctrl+F), in JavaScript bundles, in localStorage or IndexedDB—the Application tab in DevTools.
Pitfalls that are learned too late
The private API changes without warning. It has no versioning, compatibility promises, or support: the frontend team renames a field on Thursday night, and your parser collects emptiness. Protection is not a "reliable selector," but control of the structure: check that required fields are in place and of the correct type; monitor the share of empty values and the number of records in the run; skip broken records but raise an alarm if the defect exceeds 10%; store raw responses for future comparison.
The API is sometimes protected more strictly than the page. This happens regularly: HTML is served calmly, but an anti-bot check hangs on /api/, which verifies both the TLS fingerprint and the combination of headers. Then, traffic savings turn into an increase in the share of unsuccessful requests, and the gain is consumed.
Signed requests. If parameters show something like sign, hash, or _s, the frontend calculates the signature in JavaScript. Reproducing it is a separate project, and often it's cheaper to stick with HTML.
Rate limits. Private endpoints are not designed for a stream: keep 1–2 requests per second, set separate timeouts for connection and reading (for example, 5 and 30 seconds), repeat only transient errors—429, 500, 502, 503, 504—and do not touch 401 and 404. Exponential backoff with jitter is mandatory; otherwise, all workers will go for a second round simultaneously. For more details, see the analysis of timeouts and retry logic for proxies.
Legal framework. Public unauthenticated endpoints are one situation; logging into an account is fundamentally different: registration means accepting the user agreement. Personal data falls under GDPR regardless of how easily it is obtained. Facts—prices, characteristics, availability—are not protected by copyright, unlike texts and images.
When to stick with HTML
A hidden API is not always advantageous. Stick with page parsing if:
- the site is server-side and there is simply no internal API;
- the endpoint requires signing or token rotation—maintaining it is more expensive than the page;
- the API has stricter protection than public pages;
- you need the final result that the frontend assembles from multiple sources;
- you manage dozens of sites: a single HTML pipeline scales better than a zoo of private APIs with individual quirks.
What type of proxy to use for API parsing
Switching to JSON changes the calculations because the bottleneck shifts: traffic becomes low, while the requirements for IP quality and session stability increase.
- Open endpoint without authorization and without anti-bot measures. Here, data center proxies are sufficient: the data volume is small, and there's no need to pay for residential ones.
- Endpoint behind an anti-bot or tied to a session. Residential proxies with sticky sessions are needed: the token, cookie, and IP must match throughout the chain, otherwise the server will drop the session on the second request. At the same time, the bill will remain modest—gigabytes in JSON mode are consumed slowly.
- Data from a mobile application. If the web version is closed and the app delivers the same data more easily, endpoints are searched through traffic interception—this is a separate procedure discussed in the article about finding hidden API of a mobile application via mitmproxy.
In brief
Twenty minutes in DevTools often replace days of struggling with a headless browser: Fetch/XHR filter, searching for visible values, Copy as cURL—and you have a working request at hand. Then the details are resolved: transfer all headers, parse the pagination scheme, implement response validation, and soberly assess whether the endpoint is more protected than the page itself. Where the private API works, it reduces both traffic volume and the number of requests—thus lowering both the cost of proxies and the likelihood of bans.
