Category: General
The Invisible Heart of Competitor Price Tracking: How Product Matching Works
The same product looks different on every site. Here's how Senkrondata's layered product-matching approach keeps competitor price data trustworthy.
When a retailer asks "what is my competitor charging for this product?", they're really asking a much harder question underneath: "Which item in my own catalog does that listing on the competitor's site actually correspond to?"
It sounds simple. It isn't. The same pair of sneakers might be listed as "Men's Sport Shoe White Size 42" on one site and "Unisex Sneaker - White/42" on another. Barcodes are sometimes missing, sometimes mistyped, sometimes different because of packaging variants. A cosmetics item might be sold as a 3-pack on one site and as a single unit on another — and if a price comparison confuses the two, the result isn't just noisy, it's actively misleading.
The quality of competitor price tracking depends far more than people expect on this one question: can you reliably match the right product to the right product?
Why this is the hard part
Price comparison dashboards, promotion optimization, dynamic pricing decisions — all of it is built on top of matching. If the matching is wrong, everything built on it is wrong too: a price compared against the wrong item might say "the competitor is cheaper" when it's actually pointing at a different product entirely. That leads to bad pricing decisions, lost margin, or unnecessarily aggressive discounting.
Scale is what makes this hard: hundreds of retail channels, millions of products, prices changing every day. At that scale, matching by hand isn't an option — but leaving it entirely to blind automation isn't safe either, because automation's error rate compounds with scale. The answer isn't a single "magic algorithm." It's a layered approach.
No single algorithm: a layered approach
Inside Senkrondata's infrastructure, product matching runs through several layers, each operating at a different confidence level. Each layer picks up what the previous one couldn't resolve.
Layer 1 — Exact identity
When a product has a unique identifier (like a barcode/GTIN) that's present and correct on both sides, matching becomes a lookup problem: find the record carrying the same identity. This layer is fast, scalable, and highly confident — ambiguity is close to zero. Whenever it's available, it's always the first choice.
The catch is that this ideal case isn't always true: barcodes go missing, get mistyped at the source, or are simply never standardized in certain categories — apparel and fashion, with their heavy variant structure, being a common example.
Layer 2 — Smart similarity
When exact identity isn't available, the system looks at a product's other signals: name, brand, category, key attributes like color, size, or pack count. Textual similarity techniques group and score likely candidates — the goal is to narrow a large pool down to the items that are probably the same product.
This layer is deliberately tolerant by design: it's meant to run cheap and fast, shrink the candidate pool, and leave the fine-grained decision to the next layer.
Layer 3 — Contextual / AI-assisted decision
The narrowed candidate pool is handed to a layer that reasons closer to how a human would. It looks past surface text similarity and evaluates context: telling a "3-pack" apart from a single unit, treating attributes like gender, size, or color as hard constraints rather than soft signals. Using a cheap layer to eliminate broadly, then a smarter (and more expensive) layer to make the final call — that two-stage funnel is the core logic here.
Layer 4 — Expert review
For cases automation can't confidently resolve, or for particularly critical products, a match gets reviewed and approved by a person. This layer isn't a fallback bolted on as an afterthought — it's a deliberate safety net: wherever automation is uncertain, a human steps in.
These four layers aren't mutually exclusive — a single match can be supported by more than one signal at once (for example, both a rule-based signal and a degree of automation). What matters is that which layer produced which match stays traceable — so when something looks off, the question "what was this match actually based on?" always has an answer.
The principles that keep it trustworthy
However smart the layers get, a few principles have to hold or the whole system stops being trustworthy:
- Variants are not "close enough." Pack count, volume, size, color — these aren't minor differences, they're different products. Tolerance here is kept close to zero.
- A missed match beats a wrong match. Not matching something you're unsure about is safer than matching it incorrectly — because a wrong match silently turns into a wrong pricing decision downstream.
- Every match has a traceable origin. Whether it came from exact identity, a rule, an AI-assisted decision, or manual review is recorded — which is essential for auditability and debugging.
Scale and freshness
Matching isn't a one-time setup you configure and forget. Catalogs change, new products appear, competitor sites restructure their listings. That's why matching runs as a continuous process, and its output feeds a separate analytics layer built for daily price comparison reporting — so reporting across millions of matches stays fast and current without putting load on the operational database.
The bottom line
What a buyer sees in competitor price tracking is usually just a table or a dashboard. But how trustworthy that table is depends entirely on the matching logic running quietly underneath it. At Senkrondata, we don't lean on a single algorithm for that — we lean on a layered approach that starts with exact identity, moves through smart similarity and AI-assisted evaluation, and falls back on expert review when it matters most. Because a correct pricing decision can only be built on a correct product match.
If you want to trust the matching logic behind your competitor price data, talk to the Senkrondata team.
Emre
Price Intelligence & Data Engineering
Emre writes about the machinery behind competitor price data: product matching, normalization, collection at scale and the analytics layer on top.
More from EmreContact Us
Leave your email address for a detailed demo or overview session, and we will get back to you shortly.
