Every arbitrageur working with Facebook Ads, TikTok Ads, or Google Ads eventually faces the question: should they pay for a ready-made spy service like AdSpy or BigSpy, or should they build their own creative scraper for competitor ads? At a volume of 100,000 ads, the difference in costs and labor becomes critical. We analyze the economics of both options using real figures.
What are spy services and why do arbitrageurs need them
Spy services are platforms that collect advertising from competitors on Facebook Ads, TikTok Ads, and other sources, compile them into a database, and allow filtering by geo, niche, vertical, launch date, and duration of display. The most popular on the market include: AdSpy, BigSpy, PowerAdSpy, Anstrex, and SocialPeta. Arbitrageurs use them to find effective creatives, assess how long an ad stays in rotation (the longer it runs, the higher the likelihood that it is profitable), and adapt their findings to their own offers.
The logic is simple: if an advertiser runs the same creative for a month, it means it is profitable. A spy service allows finding such ads without manually monitoring thousands of pages and ad accounts. Subscriptions to such services are usually charged based on the number of requests or access to the database, without tying to the number of "viewed" ads — theoretically, you can scroll through 100,000 or 500,000 cards within one tariff if the limits allow.
However, ready-made services have limitations: databases are updated with a delay, some niche geos are poorly covered, and filters do not always provide the necessary depth of selection. If you need not general trends but targeted collection on specific advertisers, domains, or keywords in large volumes — the question of self-parsing arises.
Self-collection of creatives: how it works
Self-collection of creatives involves parsing open ad libraries (for example, Facebook Ad Library, TikTok Creative Center) using scripts or ready-made no-code scrapers. Technically, the process looks like this: an automated bot accesses the ad library page, enters search parameters (geo, niche, advertiser), extracts cards with text, images, videos, and metadata (display dates, number of active ads from the advertiser), and saves everything in a table or database.
The main difficulty is that platforms limit the number of requests from a single IP address. For example, Facebook Ad Library starts serving captchas or temporarily blocks access after a certain number of requests within a short period. Therefore, for stable collection of 100,000 ads without interruptions, IP rotation is needed — which means proxies become a mandatory element of the infrastructure, not an option.
Many arbitrageurs who already have experience with anti-detect browsers (Dolphin Anty, AdsPower, Multilogin) use a similar approach for parsing: each "profile" of the collector is tied to a separate IP and a unique browser fingerprint, so the platform does not see hundreds of requests from one machine.
Cost calculation for 100,000 ads
Let's calculate using a specific example. The task is to collect 100,000 ads from Facebook Ad Library and TikTok Creative Center within a reasonable timeframe, without bans and captchas.
Option 1: Spy service. Subscriptions to AdSpy, BigSpy, or similar services are usually sold in packages with a fixed number of "search requests" or "views" per month. If the tariff allows access to the database without a strict limit on views, the cost of collecting 100,000 ads is essentially equal to the cost of a monthly subscription — regardless of how many cards you actually extract. The downside is that some of the ads you need may simply not be in the service's database, especially if it is a narrow niche or rare geo.
Option 2: Your own scraper. Here, the cost consists of three components: proxies, computational resources (server or local machine), and time for setting up the script or no-code scraper. To collect 100,000 ads through open libraries, several hundred unique IP addresses are usually required, which will be rotated to avoid hitting the platform's limits. Data center proxies are cheaper but often get filtered out if the platform aggressively bans subnets. Residential proxies with real IPs of regular users pass protection much more stably, especially when collecting data from Facebook and TikTok, which actively ban data center ranges.
In practice, for a volume of 100,000 ads, traffic is estimated at around 5–15 GB, depending on whether images and videos are downloaded in full or just text metadata. If you only collect text and links — traffic will be minimal, and the collection can fit into a budget significantly lower than a monthly subscription to a spy service. However, if media files are needed (downloading video creatives in full) — traffic consumption and, consequently, the cost of proxies increase significantly.
For such tasks, residential proxies are optimal — they provide stable access to ad libraries without frequent captchas and allow distributing requests across hundreds of unique IPs, simulating regular users.
Technical requirements for your own scraper
If you decide to collect creatives independently, even without writing code, you will need several components of infrastructure:
- Proxy pool with rotation. The larger the volume of collection, the more unique IPs need to be involved. For 100,000 ads, it is reasonable to plan for rotation every 50-100 requests per IP.
- No-code scraper or ready-made script. There are ready-made solutions for extracting data from Facebook Ad Library without programming — they work on the principle of "enter a query → get a table." For more complex tasks (for example, monitoring specific advertisers by domains), custom setup will be required.
- Anti-detect browser or emulation of unique fingerprints. If the collection is done through browser automation, each profile must have a unique fingerprint so that the platform does not link dozens of sessions into one request chain.
- Data storage. 100,000 ads with metadata and media files amount to tens of gigabytes. A table or database is needed to store the results, plus a deduplication system to avoid collecting the same ads repeatedly.
An important point — the geography of the proxies must match the geo you are analyzing. If you are studying ads for the US market, requests from IPs in Russia or Germany will yield irrelevant or truncated results, as ad libraries often localize results based on the region of the request.
Comparison table: spy service vs own scraper
| Criterion | Spy Service | Own Scraper |
|---|---|---|
| Startup Speed | Minutes — paid and use | From several days for setup |
| Depth of Selection | Limited by the service database | Full, by any parameters |
| Data Relevance | Delay in database updates | Real-time data |
| Filter Flexibility | Standard presets | Any custom conditions |
| Need for Proxies | No, that's the service's job | Mandatory, data center proxies or residential |
| Scalability for 100k+ | Limited by the tariff | Limited only by infrastructure |
| Required Skills | None, ready interface | Basic setup of no-code tools or scripts |
Step-by-step setup for self-collection
If you decide to collect creatives independently, here is a basic algorithm of actions without the need to write code:
- Identify sources. Facebook Ad Library, TikTok Creative Center, Google Ads Transparency Center — each platform has its own data structure and request limits.
- Connect a proxy pool. Set up IP rotation in the chosen scraper or anti-detect browser — specify the type of proxy (HTTP/SOCKS5), geo, and frequency of address changes.
- Set up profiles in the anti-detect browser. In Dolphin Anty, AdsPower, or GoLogin, create several profiles, each linking a separate proxy and unique fingerprint to avoid session linking.
- Set collection parameters. Specify the niche, geo, keywords, or specific advertisers you want to monitor.
- Start collecting in small batches. Begin with 1,000-5,000 ads, check the stability of operation without captchas and bans, then scale up to the desired volume of 100,000.
- Set up deduplication and storage. To avoid collecting the same cards again, add a check for unique ad ID before writing to the table.
- Analyze by display duration. The most valuable finds are ads that run longer than others: this is a sign that the creative is profitable.
When to choose a ready-made service and when to use your own scraper
A ready-made spy service is a reasonable choice if you are doing a one-time or irregular market analysis, working in popular verticals (gambling, dating, nutrition), where the service database definitely covers the necessary ads, and you do not have the resources to set up infrastructure. It is a quick start without technical difficulties — pay, log in, and filter.
Building your own scraper makes sense if you are working with narrow niches that are poorly covered in spy service databases, need to monitor specific advertisers or domains regularly, or if the volume of collection exceeds the limits of available tariffs. Additionally, your own scraper is beneficial in the long term: with regular collection of large volumes (several times a month of 100,000+ ads), the costs for proxies and infrastructure quickly pay off compared to the accumulated cost of several monthly subscriptions to services.
Many agencies and teams of arbitrageurs use a hybrid approach: a spy service for quickly finding trends and a general market overview, plus their own scraper for targeted deep monitoring of specific competitors or niches where a complete database without gaps is important.
Risks and pitfalls
When collecting creatives independently, there are several risks to be aware of in advance. First, platforms regularly change the structure of ad library pages, which can cause ready-made scripts and no-code scrapers to temporarily stop working and require updates to the collection logic.
Secondly, using low-quality data center proxies leads to frequent captchas and bans at the subnet level — if one IP from a data center range gets blacklisted by the platform, the entire pool may come under suspicion. For parsing large volumes from open libraries, it is recommended to have a reserve of proxies with the ability to quickly replace blocked addresses.
Thirdly, the legal side must be considered: collecting data from open ad libraries (which the platforms themselves publish for advertising transparency) usually does not violate rules, but mass downloading of media files followed by copying creatives without changes can create risks of copyright claims if you use someone else's materials directly, rather than as a reference for your own creative.
For arbitrageurs who also work with TikTok Ads, it is worth noting that mobile proxies often provide more stable access to mobile versions of ad libraries, as TikTok actively tracks traffic patterns specifically from data center IPs.
Conclusion
The choice between a spy service and your own scraper depends on the regularity of tasks, budget, and required data depth. At a volume of 100,000 ads, a ready-made service wins in startup speed and simplicity, while your own scraper excels in filter flexibility, data relevance, and cost-effectiveness for regular use. If you collect data infrequently and work in popular verticals, a subscription to a spy service will save you the headache of infrastructure. If collection is needed regularly and in large volumes — it is more profitable to invest in your own scraper with a quality pool of proxies.
If you plan to build your own system for monitoring competitor creatives, we recommend starting with residential proxies — they provide stable access to Facebook and TikTok ad libraries without frequent captchas. For less demanding tasks, where speed and low cost are important, you might consider mobile proxies, which are less likely to get blocked on mobile versions of advertising platforms.