~/ecommerce/woocommerce-scraper

WooCommerce Scraper — Products, Prices, Stock & Reviews

Scrape products, variations, categories and customer reviews from any WooCommerce store. Give it a bare domain and it finds the catalog — even when the store has its JSON API switched off.

general TypeScriptCheerio Global
proooxy/woocommerce-scraper — spec
categoryecommerce / general
languageTypeScript
stackTypeScript, Cheerio
marketsGlobal
outputclean, RAG-ready JSON

key features

A bare domain is enough — the Actor finds the catalog on its own

Keeps working on stores whose JSON API is disabled or firewalled

Every variation of a variable product, each with its own SKU, price and stock status

Customer reviews as separate records with rating, author and date

Prices in minor units with the store's own formatted string alongside, so arithmetic never rounds wrong

Per-store caps, so one large catalog cannot starve the others in a multi-store run

Keyword search across every store in the run, not just a full-catalog crawl

Conservative concurrency by default — most WooCommerce stores are small self-hosted WordPress sites

use cases

  • Price and assortment monitoring across independent WooCommerce retailers
  • Building a product catalog from a long tail of small stores no marketplace covers
  • Competitive research where the competitor runs WordPress rather than Shopify
  • Review mining for product sentiment and rating distribution
  • Feeding product data into comparison sites, marketplaces and shopping agents
  • Stock and availability tracking for distribution and reseller checks

input parameters

ParameterTypeRequiredDescription
domainsarrayoptionalBare domains, one per line — the simplest way to point the Actor at a store
startUrlsarrayoptionalAny WooCommerce store URL: store root, category page, tag page or a single product
maxProductsintegeroptionalHow many product records to save across the whole run
maxProductsPerStoreintegeroptionalCap for each individual store, so one large catalog cannot consume the whole budget
includeVariationsbooleanoptionalFetch every variation of variable products, each with its own SKU, price and stock
includeReviewsbooleanoptionalSave customer reviews as separate records with rating, author and date
querystringoptionalSearch term applied to every store root or domain in the run
allowHtmlFallbackbooleanoptionalAllow the HTML fallback for stores whose JSON API is disabled or firewalled
proxyobjectoptionalProxy configuration; the scraper escalates only when a store needs it

Output Example

 1{
 2  "source": {
 3    "id": "1251",
 4    "canonicalUrl": "https://www.shoprootscience.com/shop/sample-kit",
 5    "retailer": "shoprootscience.com",
 6    "language": "en",
 7    "currency": "USD"
 8  },
 9  "title": "Sample Kit",
10  "brand": "",
11  "categories": ["Samples"],
12  "price": {
13    "current": 2000,
14    "currentFormatted": "$20.00",
15    "previous": 2000,
16    "stockStatus": "InStock",
17    "stockCount": 0
18  },
19  "stats": { "rating": 4.75, "reviewCount": 4 },
20  "options": [
21    { "type": "Sample 1", "values": [
22      { "id": "Bare Facial Serum", "name": "Bare Facial Serum" },
23      { "id": "Youth Facial Serum", "name": "Youth Facial Serum" }
24    ] }
25  ],
26  "variants": [
27    { "sku": "SK-BARE", "price": { "current": 2000, "currentFormatted": "$20.00", "stockStatus": "InStock" } }
28  ]
29}

Tips

Start with one domain and a small cap. A hundred products is enough to see which data surface that store exposes and how complete the records are. Set maxProductsPerStore in multi-store runs. Without it, the first large catalog can absorb the entire budget before the other stores are touched. Turn reviews on deliberately. They are separate records and separate charges, and most price-monitoring work does not need them. Leave the HTML fallback enabled. It is what makes locked-down stores work at all; the per-run cap keeps its cost bounded.

faq

What if the store has its JSON API switched off?
That is the common case this Actor is built for. It detects the situation and still collects the catalog, so a locked-down store returns products too.
How do I know a site is actually WooCommerce?
Point the Actor at the domain and let it decide — detection is part of the run.
Are variations separate records?
No. One record is one product, with its variations nested in variants, each carrying its own SKU, price and stock status. Reviews, when enabled, are separate records.

related in ~/ecommerce

Run WooCommerce Scraper — Products, Prices, Stock & Reviews, or get a custom build

Start extracting on Apify in minutes, or hire me to build a bespoke scraper and RAG pipeline for your exact source and schema.

run on Apify get custom data