A product catalog is not a document. Handles group rows into products, variant rows hang off those handles, image rows bind to both, and an import file rewrites all three at once. The importer does exactly what the file implies. It has no view on what you intended, and the format gives you no way to tell it.
That gap between what a file implies and what its author meant is where catalog updates go wrong, and they go wrong quietly. The import reports success. The row count matches. The damage is in the records the file never meant to speak about.
These are four of the checks that stand between a generated file and a live store. Each runs on the file itself, before it is allowed to leave.
-
Duplicate handles
The handle is the product identity, and rows that share one are not separate products. They are a single product: the first row carries the product fields and every later row extends it with a variant, an image or an option value.
So a handle that repeats by accident does not create a second product. It folds two products into one, and the update lands on the whole handle. Nothing about the file looks wrong while it happens.
The checkGroup the rows by handle before anything is sent, and reject the file if a handle repeats, or if a handle in the file is not already in the store.
# A handle is the product identity. Two rows sharing one are not two # products: the later row overwrites the earlier. counts = Counter(norm(r[HANDLE]) for r in rows[1:]) dups = [h for h, c in counts.items() if c > 1] check("no duplicate Handle (a later row overwrites an earlier one)", not dups, dups[:5]) # The mirror of it. A handle that is not already in the store does not # update the product you meant, it creates a second one beside it. unknown = [h for h in counts if h not in live] check("every Handle exists in the store", not unknown, unknown[:5]) -
Blank cells in a column that is present
The unit of update is the column, not the cell. If a column is in the file, the importer reads every row as an authoritative statement about that column. A blank cell is not "no opinion here". It is "the value is nothing", and the live value is overwritten with empty.
This is how a file built by filtering a full export down to the rows you wanted to change erases fields it never intended to touch. Absence of a column means no opinion. Presence of a column with an empty cell means delete.
The checkEvery column in the file has to be one the file is deliberately setting. A blank cell inside one of those columns fails the file unless the blank is declared as an intentional clear, and a declared clear is still refused if it would wipe a value that looks valid.
# In a MERGE import an empty cell in a column that is present # overwrites the live value with nothing. Find those before sending. deletions = defaultdict(list) for r in rows[1:]: h = norm(r[HANDLE]) for i, field in update_cols: if not norm(r[i]) and live.get(h, {}).get(field): deletions[field].append((h, live[h][field])) unexpected = {f: v for f, v in deletions.items() if f not in allow_delete} check("zero deletions in a field not declared with --allow-delete", not unexpected) # A declared deletion is not a blank cheque either. If the value being # wiped still looks valid for that field, the deletion fails too. for f in allow_delete: lo, hi = NUMERIC_RANGE.get(f, (None, None)) suspicious = [(h, v) for h, v in deletions.get(f, []) if lo is not None and lo <= first_num(v) <= hi] check(f"declared deletion of {f} does not wipe a valid value", not suspicious) -
Command columns
Catalog updates are driven by command columns, and the command decides whether a section is merged into what already exists or replaces it outright. The same rows carrying the wrong command are the difference between adding a variant and replacing the variant set.
A command is easy to inherit by accident, because an update file is usually built from an export that already had one.
The checkAssert that every command column is present, populated, and set to the operation this file is actually for. A file cannot inherit a command from the export it was built out of.
# The command decides whether a section is merged or replaced. An update # file inherits it from the export it was built out of, so assert it. if "Command" in header: ci = header.index("Command") bad = {norm(r[ci]) for r in rows[1:]} - {"MERGE"} check("Command = MERGE on every row", not bad, bad) # And the baseline the file is judged against has to be the store as it # is now: the last export plus every delta imported since. Validating # against a stale export hides deletions of values a previous delta added. live = load_live(args.live, args.live_delta) -
Image paths on variant updates
Images bind to variants by URL. A path that is wrong, unreachable or carrying a stale query string does not raise an error at import time. The media simply does not attach, and the variant goes live without an image.
This one is only visible on the storefront, which is the worst place to find it.
The checkResolve every image URL in the file before the import runs, and fail the file on any that does not come back as an image.
def head(url, retries=2, timeout=15): for _ in range(retries): try: req = urllib.request.Request( url, method="HEAD", headers={"User-Agent": "Mozilla/5.0"}) with urllib.request.urlopen(req, timeout=timeout) as r: return url, r.status except urllib.error.HTTPError as e: return url, e.code except Exception as e: last = type(e).__name__ return url, last # A bare file name is not a URL. It fails silently at import time. bare = [r for r in rows if not str(r["Image Src"]).lower().startswith(("http://", "https://"))] check("every Image Src is an http(s) URL", not bare) # Then probe every distinct URL in parallel, before the import runs. with cf.ThreadPoolExecutor(max_workers=16) as ex: status = dict(ex.map(head, unique_urls)) check("every Image Src returns HTTP 200", all(v == 200 for v in status.values()))
None of these four were designed up front. Each was written on the day its failure mode was understood, and each has run on every file since.
The build is not finished, and it is not meant to be. Every catalog that goes through it teaches it something, and what it learns becomes another check. That is the whole difference between this and a standard import pipeline: a standard pipeline handles the cases in its specification, and this one accumulates the edge cases it has actually met, then prevents them instead of reporting them.
I build the system from nothing. What makes it worth having is everything it has been taught since.