Quick answer: For an authorized catalog page, select one product record first, then extract its title, current price and currency within that record. Require one stable product ID and one unambiguous value for each required field. A selector that returns a string is only the beginning: validate whether that string is the current price for the right product, not an old price, a placeholder or a value from a neighboring card. Quarantine unexpected batches instead of turning a missing value into zero.
Important constraint: A transport success does not make extracted data correct. An HTML page can return 200 while a layout change causes a selector to pick the wrong element. No selector is a universal contract for a third-party page; explicit IDs or attributes are useful only if the publisher intends them to remain stable. Prefer a documented API or export when one is available for the permitted data. Sources: Playwright locators, Scrapy selectors.
This guide uses a hypothetical merchant-authorized feed with two product cards. The goal is to keep a catalog dataset accurate when ordinary page markup changes. It does not cover gaining access to restricted pages or collecting personal data. The method is a combination of record scope, field meaning, and batch-level alarms. Each layer catches a different failure: scoping prevents cross-card values, field checks reject malformed records, and batch checks reveal a layout that changed across the page.
Start from the record, not the price span

Suppose a page contains two cards with stable product IDs, titles, prices and currencies. The first product may also show a crossed-out old price. A global selector such as “first element with class price” can produce a perfectly valid string from the wrong card or the old-price element. Begin by identifying the card for each product, then search inside that card for the fields you need. Keep the product ID attached to every value as it moves through parsing and validation.
For example, an authorized page might expose a card shaped like this:
<article data-product-id="SKU-101">
<h2 class="product-title">Trail Bottle</h2>
<span class="old-price">$29.00</span>
<span class="current-price" data-currency="USD">$24.00</span>
</article>
A second card might be SKU-102, “Desk Lamp,” with a current price of $48.00. Those labels and amounts are illustrative input, not values fetched from a real store. The key design is the hierarchy: find the card, read the ID, then extract one title and one specifically marked current price from that same card. If the publisher offers an explicit product-data export or API, use that contract in preference to fragile rendered markup.
Playwright’s locator guidance favors user-facing attributes or explicit contracts over deep CSS or XPath chains tied to every wrapper element. Role and test-ID locators can be useful in an application you control, but a third-party catalog does not promise to retain a test ID for your collector. Choose the most stable per-record anchor actually provided and document why it is expected to remain. A long chain such as div:nth-child(4) > div:nth-child(2) > span:nth-child(3) expresses page layout rather than product meaning.
Keep nested selectors truly relative
In Scrapy, CSS and XPath can both be used to select fields. Its selector documentation distinguishes .get() for a single result and .getall() for a list. That distinction matters: calling a single-result method can hide a duplicate match, while expecting a list without checking its size can let an empty value pass unnoticed. For a required field, inspect cardinality explicitly: exactly one card ID, one title and one current price per product card unless the page contract says otherwise.
XPath has a particularly easy scoping trap. When operating on an already selected card, a nested expression starting with // searches from the document root; .// stays relative to the current card. A loop that looks correct can therefore assign the first page price to every product. A relative CSS selector can have a similar bug if it is accidentally run against the full response rather than the card object. Verify scope with a two-product fixture where the titles and prices differ visibly.
Do not rely on an exact multi-class string merely because it matched one snapshot. Sites can reorder class names or add a temporary class. Prefer an explicit field marker, schema-backed export or a clearly scoped semantic relationship where possible. When markup is ambiguous, flag the record for review rather than adding a permissive fallback that silently selects whichever span happens to come first.
Test cardinality on both an ordinary two-card fixture and a deliberately altered one. Remove the current-price element from one card; the parser should report missing, not inherit the other card’s price. Add a second current-price element; the parser should report ambiguous, not choose the first without a rule. Swap the cards’ order; the ID-to-title-to-price association should remain the same. These checks do not need a live merchant page to expose a scoping bug, and they make the selector’s assumptions visible to the next maintainer.
Validate meaning after extraction
The parser should preserve raw text and context alongside normalized values. A string $24.00 is not enough if the currency comes from a page-level setting, an attribute or a visible symbol used by several currencies. A localized format such as 1.234,50 can have different decimal and grouping conventions from 1,234.50; apply the correct locale before converting to a number. A blank field is missing, not a price of zero. A genuine zero-price item needs explicit evidence, not a default value inserted after a failed parse.
Scroll horizontally to read all columns.
| Check | Bad observation | Response before publishing data |
|---|---|---|
| Product identity | Two cards produce the same required ID | Quarantine duplicates; inspect card scope and source data. |
| Required title or price | Blank, absent or multiple current-price matches | Mark record invalid; do not substitute an adjacent card or zero. |
| Currency | Missing or inconsistent with the intended catalog context | Retain raw value and review currency source. |
| Old versus current price | Extracted value matches crossed-out old price while current value differs | Fix field selection; verify with fixture and page sample. |
| Numeric parse | Locale conversion yields an implausible magnitude | Check separator rules and original text before acceptance. |
| Batch row count | Sudden large drop or jump relative to expected catalog range | Hold the batch and inspect layout, pagination or real inventory change. |
The row-count threshold should be chosen from the feed’s normal range and business context, not copied as a universal percentage. A sale, category change or seasonal removal can legitimately change the number of products. The alarm tells an owner to distinguish a business event from extraction drift. Similarly, a price change is normal; a price that appears to move from $24.00 to $2,400.00 after a locale parser change deserves confirmation before it reaches downstream decisions.
Make accepted-record criteria explicit. For example: one unique ID, one nonblank title, one current price with a known currency, and no unresolved old-price ambiguity. An HTTP 200 and a nonempty string do not satisfy those criteria. Keep a per-record rejection reason so an operator can tell whether the page was empty, the selector matched twice, or a normalization rule failed.
Detect drift with fixtures and canaries
Save permitted, representative fixtures: one ordinary card, one card with an old price, one localized price and one missing-field case. These are review inputs, not a claim that a particular selector is permanently safe. When the parser changes, run it against the fixtures and compare the output to the expected IDs, titles, prices and currencies. A fixture that only proves “some text was returned” is too weak; it should catch a shift from current to old price or from one product card to another.
Use a small set of canary records that can be checked against the authorized source or a documented export. If the current page changes legitimately, update the expected canary after review. If every canary suddenly disappears, stop the batch before it replaces the dataset. Do not continuously hammer a page to improve a failing selector; extraction correctness and request pacing are separate controls.
During an actual run, compute validation counts before publishing: total cards found, unique IDs, accepted records, each rejection reason and any surprising row-count change. Compare them with recent accepted runs and the feed’s known structure. A new layout may preserve the same row count while swapping current and old prices, so both field-level and batch-level checks are necessary.
If a batch fails its rules, quarantine it with the raw authorized sample, parser version, timestamp and failure reason. Keep the last known good dataset available only with its honest as-of timestamp; do not present stale data as if it were just refreshed. Alert a named owner to inspect the source and fix the selector or validation rule. Once a corrected parser passes fixtures and a fresh sample, rerun the affected batch and record which published data version it replaced.
Distinguish a page change from a selector defect
A product really can be removed, renamed or repriced. The collector’s job is to avoid inventing values while it determines which happened. Check the page or official export for the affected ID. If the ID is gone and the merchant confirms removal, treat that as a business change. If the ID is still present but your parser returns no title, inspect markup and scope. If the old and current prices both exist, ensure the field rule identifies the intended current price rather than silently choosing the first match.
Keep a short change record: source sample, previous and new selector, fixture results, accepted-record counts and the person who approved resuming publication. This makes a future incident diagnosable without claiming a zero-error pipeline. The scraping hub covers other collection choices; for selectors, the practical standard is simple: no ambiguous record enters the dataset as if it were verified.
Sources and checking
Product terms can change. These are the sources checked for this article; follow the links to verify current details before you buy.
- Playwright: Locators (checked 2026-10-03)
- Scrapy: Selectors (checked 2026-10-03)