Quick answer: Give an authorized read-only collector one shared attempt budget, one per-host pacing rule and a cache that respects HTTP response directives. Retry temporary failures only within that budget. Honor a server’s Retry-After instead of sending the next request when your local timer expires; if the requested wait exceeds the job’s remaining time, defer the job. Reuse an unchanged representation only when its cache identity and validation match the new request.

Important constraint: An HTTP 200 is not necessarily an accepted data record. A login page, challenge, stale body or error message can arrive as 200 and should not trigger an endless “try again” loop. A 429 is a rate-limit signal and its response must not be cached. A 304 is useful only when the collector already holds the matching response it validates. Sources: RFC 6585, RFC 9111.

Before scheduling requests, confirm permission to read each target, its terms, and the intended data use. Set request pace and cache retention for that target and your agreement with it. The bounded settings below provide a starting policy to adapt to those requirements.

Set four separate controls

It helps to separate terms that are often lumped together as “throttling.” A timeout bounds how long one attempt can wait for a response. A concurrency limit bounds how many requests are in flight, though a target setting may not be an exact global cap. Pacing leaves time between requests to a host or service. Backoff increases or follows a wait after a retryable problem. A circuit breaker stops scheduling new work for a target after repeated failure or a rate-limit instruction. Changing only one control can leave the others free to create a burst.

For example, reducing a timeout from 30 seconds to five can make attempts fail faster and therefore increase request volume if every failure is retried immediately. Adding backoff without a shared attempt budget still allows an outer worker and an HTTP client to multiply attempts. A cache may reduce unchanged downloads, but it does not excuse exceeding an access limit or repeatedly validating every item at once. Write down the limits together, including who owns them when jobs run in parallel.

Scroll horizontally to read all columns.

ControlDecide explicitlyRecord in a request log
Attempt budgetMaximum total attempts for one eligible read, including nested clientsOriginal job ID and attempt number
DelayLocal backoff, server Retry-After, shared per-host pauseSelected wait and source of the wait
Pacing/concurrencyIn-flight requests and spacing for the authorized targetHost, start time and active count
DeadlineWhen work must be deferred rather than retriedRemaining job time and defer reason
CacheWhich representations may be stored, validated and reusedCache key, directives, validator and hit/304 status

Scrapy’s AutoThrottle is one implementation example: it adapts delays using observed latency and respects configured concurrency and minimum-delay settings. Its target concurrency is a target rather than an exact enforced ceiling. A framework feature cannot replace an explicit attempt budget or a response-specific stop rule, and a different stack needs its own equivalent controls.

Classify the response before retrying

The collector should validate the body and record, not just the HTTP status. Decide which fields make a record usable and which response means access has changed. The matrix below is a policy for a read-only GET/HEAD workflow, with a finite deadline and authorization already established.

Scroll horizontally to read all columns.

ObservationRetry decisionCache or stop action
200 with expected contentNo retry for this readValidate fields, accept the record and apply the response’s cache policy.
200 with login, error or challenge contentNo blind loopReject it as data; investigate access, parser or session state.
304 with matching stored representationNo body retry neededReuse and update the eligible cached response under HTTP validation rules.
304 without a matching stored responseDo not pretend the body existsInvestigate the cache mapping; make a normal authorized fetch within budget if appropriate.
429 rate limitWait as instructed and retry only within budgetPause relevant scheduling; never cache the 429 response.
Temporary 503 or connection failureBounded retry if the read remains eligibleApply delay and deadline; lower scheduling pressure or stop after exhaustion.
Repeated 404Stop routine retriesInspect whether the URL changed or the resource was removed.
401 or 403 denialStop routine retriesVerify authorized access and configuration; do not rotate identities as a workaround.

The 404 rule says repeated deliberately: a specific target may have a transient publication race, but that exception needs evidence and its own small budget. A generic “retry all non-200” policy treats denial, removal and rate limiting as if they were short outages. It can turn one bad URL into a large volume of unproductive requests. Likewise, a 200 body with the wrong fields is a data-quality or access diagnosis, not proof that immediately sending the same request again will fix it.

For 429, RFC 6585 says the response can include Retry-After and must not be stored by a cache. The rate-limit identity and scope belong to the server; do not assume that one URL or IP address is the unit being limited. A shared pause for the relevant host or service is often more sensible than allowing each worker to retry independently, but confirm the target’s instructions rather than guessing a universal scope.

Work through a four-attempt example

Four-attempt illustrative retry schedule has three local waits: at least one, two and four seconds plus jitter, then the fourth attempt succeeds or ends the job; a server Retry-After can override.
Classify the response before retrying and keep the attempt budget finite.

Suppose the collector allows at most four attempts total for an eligible read. It selects three local base waits of 1, 2 and 4 seconds between attempts, with small nonnegative jitter to avoid synchronized retries. Those values are example settings, not HTTP requirements or a universal polite rate. A real target may require much slower pacing or a specific contract limit.

The sequence is: attempt one fails temporarily; wait at least one second plus jitter; attempt two fails; wait at least two; attempt three fails; wait at least four; attempt four either succeeds or the job stops with an exhausted-budget result. If a server provides Retry-After, follow the applicable server wait instead of the shorter local backoff. RFC 9110 allows Retry-After to be a delay in seconds or an HTTP date, so the scheduler must parse the form it receives.

Now suppose the server requests 120 seconds but the job has 30 seconds left. Cutting the server’s wait to 30 seconds and trying again would defeat the instruction. Defer the job to a later window or stop it with a clear deadline result. Preserve the original attempt count and the server wait when it resumes, according to the workflow’s retention policy; otherwise a rescheduled job can silently restart its budget and add more requests than intended.

Share the budget across retry layers. If the HTTP client tries four times, a crawler framework retries that whole call three times, and an outer queue retries the job twice, one logical read can generate 24 attempts (4 × 3 × 2) before anyone sees a final failure. Set one owner for attempts and disable or account for the other layers. Log the original job ID, target, attempt number, delay and final stop reason so the total can be reconstructed.

This example covers read-only GET or HEAD requests. RFC 9110 treats retrying non-idempotent operations differently; a POST that creates or changes data must not be automatically replayed unless the client has specific assurance that doing so is safe or that the prior attempt was not applied. Even a GET method is not permission to hammer a target. Keep authorization and pacing separate from the method’s HTTP semantics.

Reuse only the right cached representation

A cache reduces repeated transfer only when it can identify the same representation and obey the origin’s directives. Keep the method and effective URL, including meaningful query parameters, in the identity. Account for headers named by Vary and separate authenticated identities or other private context. Do not normalize query strings by deleting parameters merely because they look inconvenient; an omitted language, variant or date parameter can make two different records collide.

RFC 9111 distinguishes storage, freshness and validation. A response marked no-cache may be stored but needs successful validation before reuse. It does not mean “never save anything.” For this collector, treat no-store conservatively as do not persist the response. Do not place private authenticated content in a shared cache. Cache-control directives are operational instructions, not a promise that data is otherwise private or correct.

When a stored response has an appropriate validator, a subsequent conditional request may receive 304 Not Modified. That 304 supplies no replacement body; it updates or validates the matching stored representation. If the cache entry is missing, tied to a different user, or mismatched on relevant request dimensions, there is nothing sound to reuse. Fall back to a correctly scoped authorized fetch within the request budget rather than creating a synthetic “unchanged” record.

Do not cache a 429 and replay it as if it were a stable page. Likewise, do not let a cache of a successful HTTP 200 freeze an invalid record indefinitely. Validate required fields when accepting a record and define when freshness expires for the business use case. A product price, public notice and static policy page can have different acceptable ages; the collector should record that decision rather than infer it from a status code alone.

Know when an operator should stop the job

Log enough to distinguish load, access and data-quality problems: request identity, host, timestamp, status, validator, cache directive, selected wait, attempt count, final stop reason and whether a usable record was accepted. Redact tokens, cookies and unnecessary personal data. Count attempts, server responses, cache reuses and accepted records separately; a high response rate can hide a low usable-record rate.

Pause and review when a target starts returning 429s, repeated 401/403s, many invalid-200 bodies or a sustained 503 pattern. Confirm that the request is still authorized, that parser expectations match the current page, and that workers share the same pacing state. The next action may be to reduce or defer collection, repair a parser, or ask the target operator for an appropriate feed or schedule. Blindly increasing retries, concurrency or identity changes is not a remedy for a denied or unstable target.

After a representative authorized run, compare requested URLs with accepted records and inspect the failures by reason. Adjust the policy only with evidence from that target and your permitted workload. The scraping hub covers broader architecture questions; this retry-and-cache policy addresses the specific risk of multiplying requests or serving the wrong stale representation.

Sources and checking

Product terms can change. These are the sources checked for this article; follow the links to verify current details before you buy.