E-commerce data pipeline · 2025

Reliable daily competitor pricing, collected on its own

ShippedIndependent engineer

Competitor prices lived in a morning ritual. Open several stores, copy figures, paste a sheet. A layout change on one site broke the day. There was no retry, no isolation per store, and no record of what failed.

The constraint

Multiple storefronts, different pagination, rate limits that are not documented. One broken parser must not take down the rest of the run.

Built with

  • Python
  • BeautifulSoup
  • Selenium

How it works

A scheduler starts the job. Each storefront is fetched and parsed on its own. Records are normalized into one table. A failed source is retried or logged, not allowed to abort the others.

The calls that shaped it

Each decision with the pressure that forced it and the price it keeps costing.

  1. One job, many sources

    One storefront changing its markup used to break the whole morning. The failure had to be confined to the source that changed.

    A scheduled process pulls each storefront independently, so a selector change fails one source without taking down the rest of the run.

    The cost: Each source carries its own parser and its own failure path, so a new storefront is real work.

  2. BeautifulSoup where the HTML is stable, Selenium where it is not

    Driving a browser for every page is slow and fragile, but some storefronts only render their prices in JavaScript.

    Not every page needs a browser. The ones that do are isolated so the rest stay cheap.

    The cost: Two parsing paths, and the browser-driven ones cost more to run.

  3. Normalize, then write

    Three storefronts describe the same product differently, and a table assembled from raw scrapes is not comparable across sources.

    The morning artefact is a table the client can open. Raw markup stays in the job's output when a source needs a replay.

    The cost: A normalization step per source, and raw markup kept alongside in case a replay is needed.

A job on a schedule, each storefront parsed on its own, a table in the morning, a log when a page changes. That is the shape of a pricing or product-data pipeline.

If you price against competitors, tell me how many storefronts you watch and how often the table needs to refresh.

Where it stands

Shipped. Competitor prices arrive as a clean table each morning, with each store isolated and failures logged.

  • A daily job collects competitor prices from multiple storefronts into one table.
  • Each storefront is fetched and parsed independently, so a selector change fails one source and the rest of the run continues.
  • A failed source is logged with its raw markup kept for a replay.
  • Shipped. The table is in the client's hands each morning.

What was handed over

  1. How the job is scheduled
  2. Per-source notes and what a selector change looks like
  3. The output table and who opens it
Next project Sales data analysis A sales review the client still opens