USE CASE

Product Data Extraction and Attribute Normalisation

Product pages describe the same attribute five different ways. Automated product data extraction reads the title, the bullet list, the spec table and the description prose together, then returns one typed, unit-normalised value per field — with the source it came from attached.

PAIN POINTS

Why Product Data Extraction Breaks Before It Reaches Your PIM

  1. 01

    The Attribute Is Never Where You Expect It

    One supplier puts pack size in the spec table. The next puts it in the title. The third buries it in a sentence of marketing prose. A selector written against one page shape returns an empty column on the other two, and the gap only shows up after import.

  2. 02

    Units and Decimals Are Not Comparable

    1,6 kg and 1.6 kg and 1600 g are the same weight. 60 Nm and 1600 in-lbs are the same torque. Extracted as strings they are four different values, and every downstream filter, facet and comparison treats them as such.

  3. 03

    An LLM on Raw HTML Invents the Missing Fields

    Pointing a language model at a product page fills every column you asked for, including the ones the page never stated. The output looks complete, imports cleanly, and is wrong in exactly the places nobody checks.

  4. 04

    The Same Page Contradicts Itself

    The title says 500 g, the spec table says 0.5 kg, the description says 450 g net. All three are on the page. An extractor that takes the first match it finds picks one at random and never tells you the other two existed.

  5. 05

    Scripts Rot Faster Than You Can Rewrite Them

    Every supplier redesign silently breaks the selectors written for it. The pipeline keeps running, rows keep arriving, and the attribute columns quietly go empty until someone notices a category has lost its filters.

EUR 0.15

Per 1,000 product pages extracted — pay for what you collect, no plan and no seats

36

Ready-made marketplace and retailer endpoints, plus any public product URL you add yourself

5

Free requests on every new account — extract your own worst supplier page before you commit

HOW IT WORKS

How Automated Product Data Extraction Works

01

Point It at the Pages

Give ScrapeWise supplier catalogues, retailer listings or a list of product URLs. Any public product page qualifies — there is no supported-site list to check against first.

02

Define the Attributes You Need

Name the fields your PIM expects and the unit each should arrive in. Supplier spec labels, in any language, are mapped into your schema rather than passed through raw.

03

Receive Typed, Traceable Rows

One row per attribute, with the normalised value, the unit, the page region it came from and its status. Delivered by API, CSV or straight into your PIM on your schedule.

One Product Page, Four Attributes, Three Different Places

This is the actual sequence behind the extraction above — a German supplier listing for a cordless drill, and the four attributes a PIM needs from it.

  1. Read Every Region of the Page, Not the First Match

    Voltage appears in the spec table as "18 V". Battery capacity appears only in the title, as "2x5,0Ah". Weight appears only in a description sentence: "Gewicht ca. 1,6 kg ohne Akku". Torque appears twice, in two units. A single-pass extractor that stops at the first match it finds would return voltage and nothing else. Each region is read separately and every candidate value is kept.

  2. Normalise the Value, Preserve the Qualifier

    "2x5,0Ah" becomes a quantity of 2 and a capacity of 5.0 Ah — comma decimal resolved, multiplier split out. "ca. 1,6 kg ohne Akku" becomes 1.6 kg with the qualifier "excl. battery" retained as a separate field. Dropping that qualifier is how two products that differ by 600 g end up looking identical in a comparison table.

  3. Send Real Conflicts to Review Instead of Guessing

    Torque is stated as 60 Nm in the spec table and 1600 in-lbs in the description. Converted, those are 60 Nm and 180.8 Nm — not a unit mismatch but a genuine contradiction in the source, most likely a spec copied from a different model. Neither value is published. The attribute is returned with both candidates, both source fields, and a review status, so a person resolves it once instead of a wrong number propagating into every channel you syndicate to.

Where Product Attributes Actually Live on a Page

Product data extraction fails on coverage far more often than it fails on parsing. These are the four regions an attribute can hide in, and what each one gets wrong on its own.

The Title

Densest source of attributes on most listings, and the least structured. Pack size, colour, voltage and capacity are routinely encoded here and nowhere else — but so is marketing language, and the two are not separated by any markup. Title parsing alone over-extracts: "Pro", "Heavy Duty" and "2024 Model" look exactly like attribute tokens.

  • raw_title
  • parsed_tokens
  • brand
  • model

The Specification Table

The most reliable region when it exists, and it frequently does not. Label wording is not standardised between suppliers — "Gewicht", "Net weight", "Weight (kg)" and "Item weight" are the same field. Values carry their own units inconsistently, and rows are often merged or split relative to the field you want.

  • spec_label
  • spec_value
  • unit
  • mapped_attribute

The Description Prose

Where the qualifiers live. Net versus gross, with or without accessories, per unit versus per case — these are stated in sentences rather than fields, and they change the meaning of a number that looks unambiguous everywhere else. Ignoring this region is why two correctly extracted weights turn out not to be comparable.

  • description_text
  • extracted_claim
  • qualifier
  • source_sentence

Images and Structured Markup

JSON-LD and microdata are authoritative when published, and most retailers publish only price and availability there — not attributes. Packaging photos often carry the only statement of case quantity or net content on the entire page. Both are real sources; neither is present often enough to rely on alone.

  • jsonld_fields
  • image_urls
  • ocr_text
  • gtin

Extracted, Inferred, or Absent

The difference between a usable attribute feed and a dangerous one is whether it distinguishes these three. Most tools collapse the last two into the first, which is why the output always looks complete.

  • Extracted — Stated on the Page

    The value was read from a specific region of a specific page, and that region is recorded alongside it. Voltage 18 V from the spec table. You can trace any published attribute back to the sentence or cell that produced it.

  • Derived — Computed, and Marked As Such

    "2x5,0Ah" yields a pack quantity of 2 that no field on the page states directly. Unit conversions land here too. Derived values are usable, but they are labelled, because a normalisation rule can be wrong in a way a quoted value cannot.

  • Absent — Returned Empty, Never Invented

    The page did not state it. The field comes back null with a reason, not filled with a plausible number. An empty column you can see is a sourcing problem; a fabricated one is a catalogue you cannot trust anywhere.

What You Send, What You Receive

You supply the pages and the attribute schema you want them mapped into. You receive one typed, normalised value per attribute, with the source region and a confidence attached.

You give
  • Product pages to extractrequired

    Supplier catalogues, retailer listings, marketplace pages or your own site. By domain, by category URL or as a list of product URLs.

    supplier-a.de/katalog/werkzeugeamazon.de/dp/B08...distributor.nl/products.csv
  • Your attribute schemarequired

    The fields your PIM expects, with the unit each one should arrive in. Supplier labels are mapped into these — you never receive raw spec labels.

    voltage:V, weight:kg, battery_capacity:Ah, torque:Nm
  • Source locale and unitsoptional
    • Auto-detect
    • EU (comma decimal, metric)
    • US (period decimal, imperial)
    • UK

    Decides how "1,6" and "1.6" are read and which unit system incoming values are converted from. Auto-detect reads it per page.

  • Fill policyoptional
    • Return absent attributes as null, never inferred

    On by default. An attribute the page does not state comes back empty with a reason, rather than filled with a plausible value.

  • Refresh scheduleoptional
    • On catalogue change
    • Weekly
    • Monthly
    • One-off backfill

    How often the source pages are re-read. Re-extraction catches supplier spec corrections, which are silent and frequent.

You get

One row per attribute per product, 11 columns each

  • product_url
  • source_sku
  • attribute
  • raw_value
  • normalised_value
  • unit
  • qualifier
  • source_field
  • status
  • confidence
  • extracted_at

Delivered by REST API, CSV export or a scheduled push into your PIM. Every row carries the page region it came from, so a wrong value is traced to its source rather than argued about.

The Attribute Rows You Actually Receive

One row per attribute, not one row per product — so a conflict on torque does not block the three fields that were read cleanly. This is the drill example as it lands in your system.

source_skuattributeraw_valuenormalised_valueunitsource_fieldstatus
SUP-A-4491voltage18 V18Vspec_tableextracted
SUP-A-4491battery_capacity2x5,0Ah5.0Ahtitleextracted
SUP-A-4491battery_count2x5,0Ah2counttitlederived
SUP-A-4491weightca. 1,6 kg ohne Akku1.6kgdescriptionextracted — excl. battery
SUP-A-4491torque60 Nm | 1600 in-lbsNmspec_table, descriptionreview — conflict
SUP-A-4491chuck_sizemmabsent
BUILD VS BUY

Build the Extractor, or Buy the Attributes

Feature
In-House Scripts + LLM Pass
Scrapewise
Coverage across page shapes
Selectors written per supplier; a redesign empties the column silently
Title, spec table, description and markup read every run, per page
Missing attributes
An LLM fills them with plausible values that import cleanly
Returned null with a reason — absent is a visible state
Units and decimals
Strings; "1,6 kg" and "1600 g" stay different values
Converted into your schema's unit, locale detected per page
Contradictions on one page
First match wins; the other values are never surfaced
Both candidates returned with their source fields and a review status
Traceability
A value in a cell, with no record of where it came from
Source region recorded per attribute, every run
What you compare on cost
Engineer time per supplier, plus token spend per product, plus the cleanup
EUR 0.15 per 1,000 product pages, 5 free requests to test
BENEFITS

Product Data Extraction That Survives a Supplier Redesign

Every Region of the Page, Every Run

Every Region of the Page, Every Run

Title, specification table, description prose and structured markup are all read on every extraction. An attribute that moves from the spec table into the title between catalogue versions keeps arriving, instead of silently emptying a column.

Normalised Into Your Schema, Not Ours

Normalised Into Your Schema, Not Ours

You define the fields and the units. Supplier labels in any language are mapped into them, decimal conventions resolved per locale, and units converted — so what arrives is importable without a translation layer in between.

Absent Stays Absent

Absent Stays Absent

Attributes the page never stated come back empty with a reason. Derived values are labelled as derived. You can tell at a glance which parts of your catalogue are genuinely sourced and which need a better supplier feed.

Your Catalogue Is Only as Good as the Attributes Behind It

Run automated product data extraction across your supplier pages and see which attributes are genuinely sourced, which are derived, and which have been quietly empty the whole time.

FAQ

Frequently Asked Questions

Common questions about product data extraction, attribute normalisation and PIM delivery.

Product data extraction is the process of reading product pages and returning their contents as structured fields — title, price, availability, images and, hardest of all, product attributes such as weight, voltage, pack size and material. The extraction part is reading the page; the part that determines whether the output is usable is normalising what was read into one schema with consistent units.