The CSS selector parser breaks when the website changes its layout. The LLM parser remains intact, but it charges for each page processed. In 2026, the choice between them ceased to be a matter of preference: the prices of the models diverged by tens of times, and the same page can cost 3,000 or 20,000 tokens depending on what you send to the model. Below is a comparison of three approaches in terms of cost, reliability, and proxy traffic, calculated for 1,000 and one million pages.
In Brief: What to Choose
- Selectors (CSS/XPath) — one template for pages, large volumes, stable layout. The cost of extraction is close to zero, but support falls on the developer.
- LLM Extraction — many different sites, unstable layout, one-time tasks. You pay for tokens on each page and must validate the response.
- Hybrid — LLM writes selectors once, then the selectors operate, and the model is called only when data validation fails. This is optimal for most permanent parsers.
Comparison Criteria
We compare based on five points that truly impact the final bill and data quality:
- cost of extraction per 1,000 pages;
- behavior when the layout changes;
- accuracy and risk of fabricated values;
- speed and latency;
- proxy traffic consumption — as you will see, it is almost unaffected by the choice of method.
How Much LLM Extraction Costs: Calculating Based on Prices in October 2026
Official prices for one million tokens (input/output) on the standard rate:
- Gemini 2.5 Flash-Lite — $0.10 / $0.40;
- Gemini 3.1 Flash-Lite — $0.25 / $1.50;
- Claude Haiku 4.5 — $1 / $5;
- Gemini 3.5 Flash — $1.50 / $9.
Google and Anthropic have a Batch API with a 50% discount on input and output — for parsing where an instant response is not required, this is the first lever for savings.
The Main Variable — Not the Model, But What You Send to It
A raw HTML page typically takes 10,000 to 40,000 tokens when sent to the model. Cloudflare, launching the Markdown for Agents feature, provided an example: the same blog post weighs 16,180 tokens in HTML and 3,150 tokens in Markdown — an 80% reduction. Other measurements for news, documentation, and product cards show reductions ranging from 67% to 94%.
For calculations, let’s assume: a raw page is 20,000 tokens, cleaned to Markdown is 3,000, plus 500 tokens for instructions and schema, resulting in 300 tokens of JSON output. For 1,000 pages, we get:
| Model | Raw HTML (20.5 million input) | Markdown (3.5 million input) |
|---|---|---|
| Gemini 2.5 Flash-Lite | ≈ $2.17 | ≈ $0.47 |
| Gemini 3.1 Flash-Lite | ≈ $5.58 | ≈ $1.33 |
| Claude Haiku 4.5 | ≈ $22.00 | ≈ $5.00 |
| Gemini 3.5 Flash | ≈ $33.45 | ≈ $7.95 |
The range is 70 times between the worst and best options for the same result. Two-thirds of this difference comes from cleaning the input, not from the choice of model. For one million pages per month, this is either about $470 or over $33,000.
Another detail: Claude models starting from version 4.7 use a new tokenizer, which, according to Anthropic, provides about 30% more tokens for the same text. When comparing bills from different generations of models, keep this in mind — Haiku 4.5 operates on the old tokenizer.
Selectors: Almost Free Until the Site Changes
Executing a CSS or XPath selector on an already downloaded page costs a fraction of a millisecond of CPU time. According to ScrapingBee's guide, on a stable layout, a regular selector is about 10 times cheaper and faster than LLM extraction. In practice, the difference is even greater because the selector does not require a network request to the model's API.
The cost of selectors lies in the support:
- if the site renames a class or wraps a block in a new div — the parser silently returns empty fields;
- A/B tests show different templates to different visitors, and some pages do not get parsed;
- on 50 different sites, you maintain 50 sets of selectors.
The most dangerous scenario is not failure, but silent data corruption: the selector grabs a neighboring element, and an old price is written to the database for weeks instead of the current one.
LLM: Resilient to Layout Changes, but Capable of Fabrication
Models do not need an exact path to an element — they search for "price" by meaning. This alleviates the problem of renamed classes and different templates. However, three typical failures arise that are described by everyone who has run such parsers in production:
- fabricated values — the model "guesses" a price or article that is not on the page;
- missing fields — some data is not extracted;
- structure drift — a string instead of a number, a different key name.
Protection is essential: strict response schema, validation (e.g., Pydantic), temperature = 0, and retries on error. Zero temperature reduces variance but does not eliminate hallucinations completely. For prices and stock levels, it is reasonable to add a check to see if "the value actually appears in the page text."
Latency is also higher: the time to load the page through the proxy is added to the model's response time — from fractions of a second to several seconds. For monitoring once a day, this is not critical; for tracking drops — it is.
Hybrid: LLM Writes Selectors, Not Extracts Data
The third approach is directly supported by popular libraries. In Crawl4AI, there is a schema generation feature: the model looks at HTML samples once and returns a set of CSS/XPath selectors, after which extraction proceeds without calls to the LLM. The documentation emphasizes that this is a one-time cost, and the schema can be reused without restrictions; with several samples, the model often selects more robust selectors based on attributes instead of fragile positional ones like nth-child.
The working scheme of the hybrid:
- LLM generates selectors based on 3–5 samples of pages with the same template.
- The parser operates on the selectors, with each record checked by a validator: fields are in place, types are correct, and prices are within a reasonable range.
- If the share of invalid records exceeds a threshold (say, 2–5%), the page goes to LLM extraction, and the schema is regenerated.
- The new schema is run on a control sample and only then replaces the old one.
This way, you only pay for the model when the layout changes, not for each of the million pages.
Summary Table
| Criterion | Selectors | LLM Extraction | Hybrid |
|---|---|---|---|
| Cost of Extraction | close to zero | $0.5–33 per 1,000 pages | close to zero + one-time calls |
| Layout Changes | breaks, often silently | usually survives | self-repairs automatically |
| Risk of Fabricated Data | none (but there is "not the right element") | exists, validation needed | minimal |
| Speed | maximum | + model response on each page | like selectors |
| Many Different Sites | expensive to maintain | strong point | good, schema for each template |
| Proxy Traffic | the same — the model does not reduce downloaded bytes | ||
About Proxies: LLM Does Not Save Traffic
A common mistake in calculations is to assume that a "smart" parser is cheaper on the network. It is not: converting HTML to Markdown occurs after downloading, so the full page passes through the proxy regardless of the extraction method. The exception is sites where the owner has enabled Markdown delivery via the Accept: text/markdown header (as in Cloudflare's function), but this is the site's decision, not yours.
For scale: with an HTML weight of 200 KB without images, 1,000 pages amount to about 0.2 GB, or approximately $0.54 on residential proxies at $2.70 per GB. Compare this with the table above: sending raw HTML to Claude Haiku 4.5 will result in a bill for the model that is 40 times higher than the proxy bill, while with Markdown and Flash-Lite, they are comparable. If you render pages with a headless browser, traffic will increase several times — measurements are available in our comparison of traffic consumption of Playwright, Puppeteer, and requests for 1,000 pages.
What truly affects traffic with any method:
- do not download images, fonts, and analytics if they are not needed;
- look for an internal JSON API instead of HTML;
- avoid unnecessary retries: each ban and retry means paid bytes. More on why the price per GB is misleading can be found in the analysis of the real cost of a successful record.
For simple catalogs without strict anti-bot protection, data center proxies at $1.50 per GB are sufficient; residential proxies are needed where the hosting IPs are cut off at the entrance.
Scenario Recommendations
- Monitoring prices of one to three marketplaces, hundreds of thousands of cards. Hybrid or pure selectors with a validator. Connect LLM only for schema regeneration.
- Data collection from hundreds of heterogeneous sites (leads, vacancies, contacts). LLM extraction to Markdown through a cheap model in batch mode. Do not run without a strict schema and validation.
- One-time research on several thousand pages. LLM: for a few dollars, you save days on writing selectors.
- Data where errors cost money (prices for repricing, stock levels). Selectors or hybrid plus checking the value against the original page text.
- RAG and knowledge bases. Here, structure is not needed, but clean text: conversion to Markdown without field extraction, the model is only at the response stage.
Conclusion
LLM extraction has not replaced selectors but shifted the decision point. The cheapest option is the hybrid: the model writes and fixes selectors, rather than reading each page. If you cannot avoid using the model on every page, first clean the input to Markdown and use batch: these two steps reduce the bill by 5–10 times before you start choosing a model. And remember that proxy traffic is not dependent on the extraction method: you need to save it on what you download, not on how you parse.
