Ingestion · 2026

Collecting web data without losing the failures

LaboratoryIndependent engineer

A data collection process that restarts on every crash, shares one rate limit across every source, and mixes broken records into the same table as valid ones quickly becomes unreliable and hard to trust.

The constraint

This is a laboratory pattern with no client metrics behind it. Four demo sources ship with it, including a JavaScript-rendered page. Tests run offline against frozen HTML so a live site cannot flake the suite.

Built with

  • Python
  • Camoufox
  • SQLite
  • Docker

How it works

Sources are YAML. Fetch writes a raw cache. Parse and validate, then deduplicate. Bad records go to a dead-letter queue. Tests use frozen HTML.

The calls that shaped it

Each decision with the pressure that forced it and the price it keeps costing.

  1. YAML per source

    Rate limits, selectors, and retries differ per source. Kept in a shared module they become globals, and one source's change turns into every source's risk.

    Rate limits, selectors, and retries live next to the source, not in a shared bag of globals.

    The cost: Adding a source is a config file plus a parser, and there is no single place that lists them all.

  2. Cache the raw HTML

    A crash partway through a run should not mean fetching the corpus again, and a parser change should be testable against the pages it already saw.

    A crash does not refetch the corpus. Replay is a local problem.

    The cost: Storage grows with every run, and the cache becomes something to manage.

  3. Two-layer deduplication and a dead letter

    A broken record that reaches the clean table is worse than a missing one: the table looks complete while it is wrong.

    Duplicates and invalid records do not enter the clean table, and retries follow the server's rate-limit signal.

    The cost: Two paths through the pipeline, and a dead-letter queue somebody has to look at.

Reliability here is earned by structure: configuration lives with each source, raw HTML is cached so a crash does not refetch the corpus, and bad records go to a dead-letter queue, which keeps the clean table clean. See the e-commerce pipeline case for that pattern running as a daily job.

Where it stands

Laboratory. A reliable pattern for collecting and validating web data, ready to apply to a client's data source.

  • Four demo sources ship with the pattern, including a JavaScript-rendered page.
  • Tests run offline against frozen HTML, so a live site cannot flake the suite.
  • Invalid records go to a dead-letter queue, and the clean table only ever holds validated rows. Retries follow the server's rate-limit signal.
  • Laboratory pattern, with no client metrics behind it.

What was handed over

  1. How to add a source in YAML
  2. How to run the tests offline
  3. What the dead-letter queue looks like
Next project Process automation suite One record, three systems, no duplication