Quick answer: Consider Scrape.do when rotating proxies, rendering pages, handling retries and maintaining fetch workers consume more of your team’s time than the extraction itself. A self-built HTTP, proxy and browser stack can remain reasonable for permitted, stable, mostly simple HTML targets when you already operate those components. The decision needs a representative authorized target set and a cost-per-accepted record ledger; published features alone do not show which route is cheaper or retrieves more usable data. Sources: Scrape.do documentation, Playwright, Scrapy.
Important constraint: A charged Scrape.do response is not necessarily a usable record. Its request-cost rules count 2XX, 400, 404 and 410 responses among chargeable successes and allow some domains to have different credit profiles. Inspect the
Scrape.do-Request-Costresponse header and validate extracted data before computing a cost per accepted result.
Start with the right to collect the data. For either architecture, check authorization, the target’s terms, applicable data rights and its crawler instructions before planning volume. RFC 9309 defines the Robots Exclusion Protocol; it does not by itself grant a right to obtain or reuse a site’s content.
Define a usable result before comparing tools

A crawler can report an HTTP response while the data team still has nothing it can use. A page might be a login screen, a stale cached view, a consent prompt, a duplicate, or a successful response missing the fields required by your downstream job. Set an acceptance rule before a trial: for example, “the current product page, with title, listed price, identifier and update date present, and no duplicate identifier.” The fields are an illustrative schema; the target’s actual permissible data and required quality belong in the team’s own specification.
That rule separates three counts that should never be collapsed into one: attempts, chargeable responses, and accepted records. A retry increases attempts. A 404 can count toward Scrape.do credits under its published rules while yielding no record. A 200 can also fail your acceptance rule. On a self-built stack, even an unusable response consumes some infrastructure and operator attention. The denominator that matters for the business job is accepted records, provided both paths use the same acceptance rule.
For a fair comparison, use the same authorized target corpus, observation window, output schema and allowed access methods on each path. Record failures rather than quietly removing hard pages from one side. If the self-build path does not render JavaScript but the managed path does, label that difference as a capability choice and compare a separate rendered workload rather than blending two unlike jobs into one “success rate.”
What the managed layer changes
Scrape.do’s getting-started documentation describes a fetch API with managed proxy handling, a render=true headless Chromium path, automatic retries and an Async API option. Those capabilities can replace components a small team would otherwise wire together and keep running. They do not determine whether a particular target will respond, whether the returned page meets your data rule, or whether the target permits collection.
Scroll horizontally to read all columns.
| Work to complete | Scrape.do documented path | Self-built path and owner |
|---|---|---|
| Fetch HTML | Send a request through the managed API. | Maintain HTTP clients, target rules and request scheduling. |
| Render JavaScript | Select the documented render option when needed. | Run browsers and their compatible binaries and system dependencies. |
| Use proxies and retries | Use managed proxy and retry controls, then inspect response and charge. | Source proxies, define retry policy, rotate identities where permitted, and log outcomes. |
| Respect target load | Set request volume and check the service’s account limits. | Tune concurrency and delays; tools such as Scrapy AutoThrottle help with that control. |
| Validate data | Build extraction and acceptance checks after fetching. | Build the same extraction and acceptance checks. |
| Handle incidents | Diagnose bad responses, changed domains and vendor/account limits. | Diagnose the browser, proxy pool, queue, workers and target changes directly. |
The right-hand column is not free merely because the frameworks are available. Playwright’s browser documentation and CI guide describe browser binaries and dependencies that a rendered worker must install and keep compatible. Scrapy AutoThrottle offers adaptive delay and concurrency controls, but it does not supply a proxy service or a browser renderer. A team already running these pieces well may prefer the control; a team repeatedly repairing them may value the managed boundary.
The API path still leaves work with the customer. Someone must define valid records, parse responses, detect layout changes, store data, handle target-specific permission rules and decide whether a suspicious result is worth another charged attempt. Managed fetching shifts a subset of operations; it does not make the data correct by default.
Convert requests into credits before looking at dollars
For untargeted domains under Scrape.do’s published base-cost table, the mode changes the credit weight:
Scroll horizontally to read all columns.
| Request mode | Published base cost per chargeable response |
|---|---|
| Datacenter proxy, no rendering | 1 credit |
| Datacenter proxy with rendering | 5 credits |
| Residential or mobile proxy, no rendering | 10 credits |
| Residential or mobile proxy with rendering | 25 credits |
Those are starting weights, not a fixed cost for every domain. Scrape.do says it can apply domain-specific profiles that change the charge, and the Scrape.do-Request-Cost response header confirms the actual request cost. The shorter idea of “pay for successful requests” can hide two distinctions: some 4XX responses are chargeable under the documented rules, and one rendered response may cost several credits. Use the header values from the actual requests in a trial, not a count of calls multiplied by the lowest row.
Consider a hypothetical monthly mix of 20,000 chargeable ordinary datacenter responses and 5,000 chargeable datacenter-rendered responses on domains that use the base weights. The first group uses 20,000 × 1 = 20,000 credits; the second uses 5,000 × 5 = 25,000 credits. Together that is 45,000 base credits. It says nothing yet about how many of the 25,000 responses satisfy the data team’s acceptance rule.
Now suppose 500 of the 5,000 rendered requests need residential or mobile proxy with rendering. Their base weight becomes 25 instead of 5. Replacing 500 × 5 with 500 × 25 adds 10,000 credits, producing 55,000 base credits for the same number of chargeable responses. Domain profiles or changed settings may alter the real header total again. This sensitivity check shows why the share of JS-rendered and residential requests belongs in the workload ledger.
Scrape.do’s public pricing page displayed a Hobby offer of $29 billed monthly and 250,000 successful API credits on October 3, 2026. That public plan figure is context for checking the live offer, not a per-record price: an actual account may have a plan minimum, taxes, mode overrides, failed acceptance checks or other terms. Its free-plan wording about “1,000 successful API calls” also does not tell you how many credits a multi-credit rendering call consumes. Keep calls, credits and accepted records in separate columns before making a purchasing conclusion.
Keep a workload ledger for both paths
Create one row per target group and fetch mode. If different domains have different permission or technical conditions, keep them separate. The ledger should be detailed enough to explain a surprising bill or a failed extraction without relying on memory.
Scroll horizontally to read all columns.
| Ledger field | Managed API measurement | Self-built measurement |
|---|---|---|
| Target and authorization | Domain, permitted data, allowed request method and robots/terms review | Same target and access record |
| Demand | Planned URLs, attempts, retries, render share and proxy mode | Planned URLs, attempts, retries and browser share |
| Accepted output | Required fields, valid records, duplicates and rejected-result reason | Same acceptance test |
| Metered use | Sum of Scrape.do-Request-Cost headers and the current plan terms | Proxy fees, browser compute, worker hosting, queue, storage and monitoring |
| People and incidents | Integration, parsing, account review and vendor escalation time | Proxy, browser, queue, parsing and incident-response time |
From that ledger, the useful comparison is total relevant spend ÷ accepted records, paired with the time and risk the team is willing to own. For the API, start with the actual billed credits and the applicable plan, minimum and overage terms. For self-build, include directly paid infrastructure and explicitly priced engineering hours. A hosted worker that needs constant repair can be expensive even with low proxy fees; a simple permitted HTML job may be economical for a team that already runs the pipeline. Neither conclusion can be calculated from a vendor feature page alone.
Do not let concurrency numbers masquerade as throughput. Scrape.do’s current Async API page and an April 2026 changelog show different plan concurrency tables, while the pricing page presents another customer-facing context. The older page may be stale, but the discrepancy makes a fixed capacity comparison unsafe. Confirm current limits in the chosen plan or account, then measure accepted records over the same period. More concurrent slots do not guarantee more valid output when target latency, retries or validation dominate.
What a small trial needs to reveal
Choose a sample that represents the expected month, not merely the easiest pages. Keep simple HTML, JavaScript-dependent pages and domains with special response rules in separate groups. For each group, record the number of URLs attempted, retries, HTTP statuses, accepted records, rejected results, elapsed time and actual credited response headers. Use a rejection label such as “missing required field,” “duplicate,” “login page,” or “outdated content” so the reason for a poor result remains visible.
The same sample can be run through the self-built route where access rules allow it. Record proxy spend, browser-worker time, queues or storage, and the staff hours needed to fix or maintain the run. The useful ratios are accepted records ÷ attempts and total relevant cost ÷ accepted records. For Scrape.do, add the returned request-cost headers first; for the self-built route, add paid services and a deliberately chosen value for staff time. Then compare the paths separately for static and rendered target groups. A single blended average can hide a managed service that helps on complex pages and adds little on simple ones.
Repeat the test when a target changes its layout or access policy, because both extraction validity and the work needed to fetch a page can change. Keep the same field-validation rule when you rerun it; otherwise a better-looking “success” figure may simply reflect a looser definition of a usable record.
A decision after the trial, not before it
Scrape.do is worth examining when several authorized targets require rendering or varying proxy modes and those moving parts repeatedly absorb the team’s development time. The managed API offers documented controls for that work, and the credit header gives a unit to record. A self-built stack remains plausible when pages are stable, permissible and mostly simple HTML, and the team wants direct control over browser execution, retries and storage. A mixed architecture can also be sensible: keep straightforward targets on an existing crawler and trial the managed path only where browser or proxy operations are consuming time.
Before committing to either route, test a representative authorized sample, inspect target-valid records, record the actual credit headers and measure the staff hours spent maintaining each path. Check the selected mode’s current domain cost and the live plan terms before treating the hypothetical 45,000-credit worksheet as a budget. The same data-quality and access checks apply whichever fetch layer you choose. The scraping hub groups related decisions about collection methods and operating trade-offs.
Free plan includes 1,000 API credits. Check credit usage and paid-plan limits for your workload.
Sources and checking
Product terms can change. These are the sources checked for this article; follow the links to verify current details before you buy.
- Scrape.do: Getting Started (checked 2026-10-03)
- Scrape.do: Request Costs (checked 2026-10-03)
- Scrape.do: Pricing (checked 2026-10-03)
- Scrape.do: Async API (checked 2026-10-03)
- Scrape.do: April 2026 Async API update (checked 2026-10-03)
- Playwright: Browsers (checked 2026-10-03)
- Playwright: Continuous integration (checked 2026-10-03)
- Scrapy: AutoThrottle (checked 2026-10-03)
- IETF RFC 9309: Robots Exclusion Protocol (checked 2026-10-03)