Quick answer: For an authorized product-price snapshot, keep the few fields needed to compare prices and a separate lineage record that answers: which source response was used, when it was retrieved, which parser version transformed it, who was responsible, where the output went, and when each copy is due for review or deletion. Give every accepted snapshot and export its own ID; do not overwrite yesterday’s evidence in place. Sources: W3C PROV, W3C data version practices.

Important constraint: A hash does not prove the data is true or that collection was authorized. An HTTP ETag is not your SHA-256 digest, and Last-Modified is not your collection timestamp. Store each with its correct meaning, and keep the actual authorization decision in an accountable system. Sources: HTTP validators, SHA standard.

Consider a team permitted to collect a daily, nonpersonal snapshot of displayed list prices for a known set of product SKUs. The business question is narrow: which listed prices changed, in which currency, between two accepted snapshots? The team does not need customer reviews, seller addresses, image files, page analytics, or full descriptions to answer it. Yet the final price table still needs enough context to explain a change if a parser, source page, or currency format shifts.

This guide starts after access permission and a response have been obtained. It does not decide whether a public page can be collected, interpret a site’s terms, or cover personal data. It uses a plain relational or JSON-shaped ledger rather than requiring a provenance graph database. The concepts come from current W3C, IETF and NIST references checked October 3, 2026; the retention periods and fields in the example are choices the actual data owner must review.

Declare purpose and authorization before the first record

Give the collection a short purpose statement: “Compare changes in displayed list price for these approved SKUs once per day.” Link to an internal authorization reference such as fictional AUTH-2026-014. That record should name the owner, source scope, allowed fields, collection frequency, and review date. The ID is a pointer to a decision, not proof of legal rights by itself. If scope changes, pause and have the responsible owner review it before collecting or retaining more fields.

Then define the output: one accepted snapshot with a product ID, displayed name for human QA, listed price in minor currency units, and currency. Store the raw listed price briefly only if needed to diagnose parsing or locale conversion. If the page unexpectedly contains free-text personal information, stop that run and route it through the organization’s separate review process; it is outside the stated dataset. “Publicly visible” alone is not the same as “needed for this purpose.”

Scroll horizontally to read all columns.

Candidate business fieldDecision for the price-change purposeReason
Source product IDKeepStable join key across daily snapshots.
Displayed product nameKeepHelps a human spot a mismatched SKU extraction; not the stable key.
Listed price, raw textConditional short retentionDiagnoses decimal and currency parsing; expire with the review window.
Price in minor units and currencyKeepMakes numeric comparison interpretable.
Description and image URL/bytesDropNot needed for price comparison.
Reviews, seller contact, analytics IDsDropUnrelated and may carry identifying or unstable detail.

Keep provenance fields—source locator, retrieval event, response status, parser version, authorization reference—outside the product business-field list. They explain where the dataset came from but are not facts about the product. A field-necessity review should challenge both sets: a URL may be useful to trace the source yet should not include a secret token in a query string; a raw response may be useful for a short debugging window but unnecessary for the longer normalized dataset.

Give every stage an ID and a responsible activity

An authorized response R417 is used in extraction X417, then normalization N417, producing accepted snapshot S417; each stage keeps its own identity and responsible activity.
Record what produced each snapshot so a parser change does not silently rewrite earlier evidence.

The W3C PROV primer distinguishes an entity (a thing such as a response or dataset), an activity (a collection or transform that uses and generates entities), and an agent responsible for an activity. Apply that model lightly. One collection run creates a response entity. An extraction activity uses it and generates an extracted entity. A normalization activity uses that and generates a price snapshot. An export activity uses an accepted snapshot and generates a delivered file. A service account or named team role is associated with each activity.

response R417 → extraction X417 → extracted E417 → normalization N417 → snapshot S417 → export P417

The arrows mean “used to produce,” not merely “happened near the same time.” W3C’s data guidance recommends recording dataset origins and changes and giving versions suitable identifiers. Keep immediate source-to-output links so a parser change is visible. If snapshot S418 uses a new parser, create a new activity and output entity even if it comes from the same URL. Do not rewrite S417 and lose the ability to explain yesterday’s exported price.

Scroll horizontally to read all columns.

Record typeMinimal useful fieldsExample
EntityID, type, source locator, schema version, digest/input rule, status, retention ownerS417, normalized-price/1, accepted
ActivityID, input IDs, output IDs, code/config version, start/end time, outcomeN417 used E417, generated S417
AgentService or accountable role ID and responsibilitycatalog-collector, data owner
DependencyProducer entity, consumer/export ID, delivered time, retention ownerS417 supplied report P417

These tables can be ordinary database rows. PROV-O provides a formal ontology if interoperability calls for it, but a small team can preserve the same relationships without adopting RDF. What matters is that an exported price can be traced back to one source response and one version of the code and configuration that interpreted it.

Keep the timestamps and byte checks honest

Record retrieved_at and transformation times in an RFC 3339 form with an explicit UTC offset, such as 2026-10-03T04:00:00Z. If the source page asserts a publication or update time, store it separately as source_asserted_at. A collector timestamp comes from your system clock; a publisher date comes from the source. Neither should silently stand in for the other. For a scheduled daily run, a failed attempt and a later successful attempt need distinct event IDs and times.

Keep HTTP response metadata with its own labels. RFC 9110 defines ETag as an opaque validator for a representation, with strong and weak forms, and Last-Modified as origin-provided modification information. Preserve the exact header values when useful. Do not treat an ETag as a local content digest or assume Last-Modified is the moment your collector saw the page. Both can be absent.

If you compute a SHA-256 digest, name which bytes were hashed: the exact stored response body, the extracted JSON encoded under a named canonicalization procedure, or the final export file. The same algorithm applied to differently encoded or reordered content can produce different digests. A stable digest can show that bytes match a prior recorded set under the same procedure; it does not prove who served the content, whether it was accurate, when it was collected, or whether the collection was permitted. Those questions require other records and controls.

Keep failed runs out of the accepted series

Suppose the collector receives a product price but no currency, or two rows claim the same source product ID with different prices. Do not let that run overwrite yesterday’s accepted snapshot. Create a quarantined entity with the response ID, parser version, error class, and owner for investigation. The next accepted version should be an explicit new entity after the issue is resolved. Downstream reports should select the most recent accepted snapshot, not the newest row regardless of status.

Use acceptance checks suited to the purpose: every retained row has a product ID, normalized price, and currency; duplicate IDs are resolved or quarantined; a sampled name matches the intended product; totals and currency changes are reviewed; and the export resolves to its source and transform IDs. If a parser update changes output from the same stored bytes, record the new parser activity and compare the two outputs. The difference might be an intentional fix or a new bug, but it must not be invisible.

The lineage record is also a practical correction route. If a product’s displayed price is wrong in a report, locate the export, its snapshot, the normalization activity and raw/extracted input if retained. If the raw response is gone under the retention policy, the remaining digest and activity record can still identify which version was used, but they cannot recreate content that was never kept. Decide the needed debugging window explicitly rather than promising perfect reconstruction forever.

Give each copy a review and deletion trigger

Retention should follow the declared purpose and the organization’s actual obligations. As an illustrative worksheet only, a team might hold an authorized raw response for seven days to troubleshoot parsing, review normalized price snapshots after 90 days, and review provenance/activity records after 180 days. Those numbers are not universal defaults. A different purpose, source agreement, or downstream contract can require different decisions. Each entity class needs an owner, a review date, a delete trigger, and a rule for open investigations.

Scroll horizontally to read all columns.

Entity classExample decision to replace with your policyWhat must happen at review
Raw responseShort troubleshooting window, if permittedDelete or extend for a documented issue; check debug copies.
Quarantined extractionUntil investigation closes, with a bounded dateResolve, reject, or delete; never promote silently.
Accepted price snapshotRetain while price-change history serves the purposeCheck whether older versions and fields are still needed.
Provenance rowKeep while an accepted output or obligation depends on itRetain only the trace needed for remaining entities.
Export or downstream copyRecipient-specific date and ownerConfirm use, request deletion when due, record response.

When an authorization changes or the purpose ends, identify every dependent report, file, API feed, or partner export before deleting only the collection table. Record which copy received which snapshot ID and who owns it. A deletion job can mark expires_at, act on known storage locations, and write a minimal event with entity ID, time, agent, outcome, and any retry owner. Do not put the deleted payload into the deletion log. A recorded acknowledgement is operational evidence, not a guarantee that every unregistered copy has disappeared.

Use a reusable blank check before accepting each new dataset version:

Scroll horizontally to read all columns.

Question for the ownerFill in before publication or export
What decision does this dataset serve?Purpose, approved fields, owner, review date: _____
Why may this source be collected?Authorization reference, source scope, frequency: _____
What produced this snapshot?Response ID, retrieval time, parser/config versions: _____
Which fields are needed?Keep/drop decision and reason for each: _____
Where will it go?Consumer IDs, delivery time, retention owners: _____
When does each copy end?Review/delete trigger and dependency action: _____

If the purpose or authorization reference is missing, a product field has no necessity reason, an export cannot resolve back to a source snapshot and transform, or a retained copy has no owner and review trigger, hold the dataset. A traceable, smaller price series is easier to correct and retire than an unexplained archive of entire pages. The same discipline remains useful when collection technology changes: keep the reason, source, transformation and lifecycle visible with every version.

Sources and checking

Product terms can change. These are the sources checked for this article; follow the links to verify current details before you buy.