Web data collection

Start with a permitted source and the fields you need. Then choose how to fetch pages, verify extracted values and handle retries. A managed API can reduce infrastructure work, but it cannot decide whether the source may be used or whether a returned value means the right thing.

API or your own stack?

Price a managed API and your own proxy/browser stack against the same authorized pages and required rendering steps. The self-managed route adds browser and proxy upkeep; the managed route still needs budgets and output checks. A successful HTTP response does not prove the intended record was collected.

  • Count billed units and operator work for one fixed page set.

Extract values that mean the right thing

A selector may still return text after a layout change while switching from the current price to a crossed-out or related-item price. Scope it to the intended item, then validate identity, result count, currency and availability. Quarantine changed layouts rather than loading plausible but wrong values.

  • Keep ordinary, missing and changed-layout examples in the validator fixture.

Retries, pacing and caching

Retry timeouts and transient server errors only within a finite budget; stop on denials, authentication failures and missing pages. Respect Retry-After and applicable cache instructions. Log attempts as well as final outcomes so a schedule does not create more load without improving the dataset.

  • Set retry ceilings, pace and cache retention for each permitted source.

Keep the dataset traceable

A price snapshot needs its source page, capture time, extractor version and validation result beside the value. Keep enough evidence to investigate a disputed record without retaining unrelated page content. Plan how corrections and deletions reach derived tables and exports.

  • Trace one exported value back to its capture and extractor version.