Alfa alfa · creating something from nothing Tal Weiss {{ roleLabel }}
{{ availabilityLabel }}
AI Solutions Architect · eCommerce catalog and data systems

Catalog systems built, verified, localized, and kept current

Turning raw product data into launch-ready catalogs: structured, enriched, localized for each market, and verified. Then keep them current, add new products, close data gaps, and sync updates across ERP, POS and storefront.

Get in touch See the work See a live store Download CV (PDF)
Shopify Magento Mirakl WooCommerce Rithum / DSCO Odoo ERP API · Feed · EDI Safe Delta Updates (Zero Overwrites) Reconciliation ERP / API sync Re-run Automation LOCALIZE
Remote / Hybrid · EU · Israel · Spain +972 54-669-9747 Talw1982@gmail.com linkedin.com/in/tal-weiss
The core pipeline
ONE CORE · ONE ADAPTER PER BRAND
RECON SIZED
Retail data flows & Data Reconciliation Source discovery Volume count Variant model Field gap analysis Access & API check Effort estimate
↓
BUILD DATA + FEATURES
Translation pipeline Feature icons Spec mapping Rating meters Sitemap & categories Smart filters Internal linking Image mapping per variant & color
↓
LOCALIZE SUB-PIPELINE

Optional in theory, on every current project in practice: international brands landing in Hebrew.

Locked glossary RTL handling Spec-language translation Terminology QA Editorial decisions recorded
↓
VERIFY ANOMALY HUNT
Missing-field detection Duplicate checks Variant integrity Broken image sweep Terminology QA Corrections as rules
↓
HANDOVER EXPORTABLE
Platform import files ERP / API sync Re-run in waves SOP & audit trail
SCOPE · I build the structure the storefront runs on: the mapping and the data model are matched to the platform they have to run on, so the catalog fits the programming rather than the other way round.
00 / The object

A catalog is a structure, not a spreadsheet

Every cube is one variant row from the latest pipeline run. Individually they are nothing. Held in the right shape, mapped, translated and verified, they are a store that opens.

11,082
variant rows in the shell
1 core
at the center, one adapter per brand
12
agents in the build pipeline: scrape, translate, map, verify
Drag to spin it
6
builds run in parallel, across four clients
six builds carried at once, each with its own source and its own platform
108,999
SKUs built
15,134
product pages built and delivered
11,082
variant rows, latest pipeline run
3 days - 2 wks
to first delivery, depending on the source
54
checks in the certification suite, plus 13 on update safety
01 / The work

What is shipping right now

Six builds across four clients, counted rather than remembered. Every number comes from a file. First delivery takes three days to two weeks, depending on how much usable source data already exists. What happens after that is not on my side: development, design, and above all the missing data and the answers only the client can give. Every plan lists what I am waiting on and who owns it.

Updated September 2026

Brand names are published only after a store goes live. Anything still in build, staging, or review appears under its code name. The live sports brand is named because its store is public. The rest follow as each one launches.

#ProjectBanana BUILT · IN HANDOVER · LAUNCH IMMINENT FLAGSHIP

Jewelry: a catalog that did not exist

Source → target
Manual work files → Odoo ERP → Shopify
2 wksto first delivery, then: client descriptions, ERP readiness Direct APIERP link, both directions Autoexport to POS and storefront
RECON ✓ BUILD ✓ VERIFY ✓ HANDOVER · NOW
The problem

The information did not exist. Not messy, absent. A register snapshot that no longer matched reality, scattered sheets, product photographs with no data attached. By any current standard the catalog was analogue.

41,178SKUs written 867models 1,586pages
What was built

Not a cleanup. A rebuild. The information structures themselves had to be designed before anything could be filled.

41,178 SKUs across 867 models, written from product imagery through vision models at scale, SKU scheme designed and not inherited.

Category architecture, collection logic, attribute filters and internal linking built alongside the data.

Direct API link to an Odoo-based ERP, auto-export to the POS and the storefront.

Unlocked supplier choice and ERP readiness, two commercial capabilities they did not have.

Case study PDF →
#GetYourRunningShoes 1 LIVE · 1 QA · 1 BUILD · 1 RECON

Sports distributor: four brands

Source → target
Global brand PIM / DAM → API collection + mapping → Local Shopify store
67,821SKUs verified across three brands 4adapters, one core 2 wksto first delivery, then: per-brand source access, launch drops

Four international sports brands under one local distributor. Source data sits in each brand's own PIM or DAM, collected and mapped through their APIs, and every brand means something different by "in stock". Not a data lift: a local Shopify store and a full product experience per brand.

LIVE Brand 1 · Saucony Israel · 156 products on site, built in full to 219 models / 701 products / 7,136 SKUs. Launching in drops, drop one shipped.
QA Brand 2 · 48,438 SKUs across 1,715 products. Advanced QA, final designs in, live catalog behind a password.
BUILD Brand 3 · recon and source collection complete at 1,148 products / 12,247 SKUs, in production.
RECON Brand 4 · recon complete at 3,253 products / 1,357 models, build queued behind brand 3.
Beyond the data
Image mapping per variant & color Full Hebrew translation pipeline Feature icon sets Technical spec extraction Sitemap & category tree Smart filters Internal linking
CushioningLevel 4 / 5
FlexibilityLevel 2 / 5

Rating scales derived per model from manufacturer spec language, then rendered as a comparable measure across the whole catalog.

Visit the live store
#ProjectWhatsTheTime BUILT · TECHNICAL REWRITE

Watches: twelve luxury brands, one catalog

Source → target
12 luxury brand sites → Master catalog + glossary → Shopify storefront
2 wksto first delivery, then: glossary decisions, domain cutover 4,344product pages
12 BRANDS
RECON ✓ BUILD ✓ VERIFY · NOW HANDOVER

Twelve luxury watch brands, each publishing a different data shape, different bot defenses, and a vocabulary that barely exists in Hebrew.

4,344 product pages unified into one master catalog.

Scraping infrastructure surviving Akamai and Cloudflare, multi-process with resume.

Hebrew haute-horlogerie glossary built term by term, with recorded editorial decisions.

Movement, complication and material taxonomies driving the filters, not just the text.

The core absorbed brands nine to twelve as adapters, not a rewrite, including a dedicated ultra-luxury sub-store.

Terminology with no Hebrew equivalent

Haute horlogerie vocabulary that does not exist in Hebrew, or exists as an insult. Each term is a decision: keep the French, keep the finish name, or coin something that a collector will accept. Machine translation gets every one of these wrong.

mother-of-pearl → אם הפנינה (Em HaPnina)
hairspring → קפיץ איזון (Balance spring)
power reserve → עתודת כוח (Power reserve)
Côtes de Genève → Preserved in French (Editorial rule)
flinqué / barleycorn → Finish name preserved (Guilloché spec)
openworked → משוחררת (Ajouré / Openworked)
oscillating weight → משקולת מתנדנדת (Rotor weight)
monocrystalline silicon balance-spring → קפיץ איזון סיליקון (Silicon spring)
complication → מנגנון מורכב (Complication)
skeleton → מנגנון שלד גלוי (Squelette / Skeleton)
#ProjectAtelier TEMPLATE APPROVED · IMPORT STAGED

Designer boutique: rebuild, not a lift

Source → target
Legacy storefront → Rebuilt catalog + URL map → New Shopify store
6,788product pages 26,275images, zero broken 3 daysto first delivery, then: design sign-off, client runs the import
RECON ✓ BUILD ✓ VERIFY ✓ IMPORT · CLIENT RUN

Not a migration. The base data had to move, the storefront presentation had to be rebuilt from nothing, and the gaps in between filled before any of it was sellable.

6,788 product pages from 2,149 models, on a per-color model the platform does not natively express.

26,275 images downloaded, zero broken, re-hosted on the platform CDN.

Technical specs written where no source existed. SEO completed per product, not inherited.

Category tree, filter attributes and old-to-new page mapping so rankings and traffic survive the move.

Web & delivery LIVE AND MAINTAINED

Sites built, run and kept current

~95live pages to date 66sections mapped 0release cycles

A practice needs a face, and so does a body of work that has none. Both have to stay current without a redesign every time something changes. Four live sites, each one visitable.

LIVE PORTFOLIO My own brand site Updated on demand through an automated flow. No release cycle, no agency, no CMS. talweiss.netlify.app → LIVE VOLUNTEER Local dive club My community club, done unpaid. WordPress → static: 26 pages, 563 assets, 66 sections mapped, guarded publish agent. indigo-club.co.il →
LIVE CLIENT Artist portfolio Six pages with a full works catalog, built to be run by a non-technical owner. Shown on request
GATED PRIVATE Submissions site Password-gated, per-page encryption. Walkthrough on request. Link on request
04 / Capability layer

It is not only the data

Every build carries the same layer inside the BUILD phase. Nine moving parts that turn a spreadsheet into a store people can actually shop, the same nine on every project.

01
01

Retail data flows & Reconciliation

Automated compatibility check matching supplier feeds and databases against target platform schemas to determine build complexity.

02
02

Full translation pipeline

Hebrew and multilingual content at catalog scale, with locked glossaries, RTL handling, and recorded terminology decisions instead of machine guesses. Grounded generation against a controlled vocabulary, not free translation.

03
03 · CORE CAPABILITY
Enterprise DAM + Vision

Image mapping, Enterprise DAM & Vector Visual Search

High-throughput ingestion connecting directly to enterprise DAM systems (Adobe Dynamic Media / Scene7, Bynder, and media hubs). Thousands of master assets matched to the exact variant and colorway, angle-tagged, deduplicated via perceptual hashing (pHash), re-hosted on CDN, and verified so zero broken links ship.

Vector Visual Search · Beta AI Vision Embeddings

Deep visual feature embeddings for automated colorway verification and metal/material recognition (yellow/white/rose gold, stainless steel, ceramic, matte finishes) directly from raw product imagery.

Read Technical DAM Case Study → 3 architectures cracked · 18,000+ assets
04
04

Feature mapping & icons

Manufacturer feature language normalized into a fixed set, then matched to icon sets so the same feature reads the same way on every product page.

05
05

Specs & rating scales

Specs extracted, written where no source exists, and expressed as comparable measures: cushioning, flexibility, support, drop. Numbers a shopper can compare across models. Constrained synthesis with provenance rules: blank if there is no real source, never an invented value.

06
06

Sitemap & category architecture

The whole site mapped: category tree, collection logic, landing structure, and internal linking designed so the catalog is navigable and indexable.

07
07

Smart filters

Filters built from real attribute data rather than free text, so size, width, terrain, movement or material actually narrow the catalog instead of emptying it.

08
08

Verification gate

Missing-field detection, duplicate and variant checks, anomaly hunting, and corrections recorded as rules. Nothing leaves without an audit trail. An evaluation harness for catalog data: fifty-four checks across nine categories, plus update safety.

09
09

Multi-agent storefront sweep

Multi-agent orchestration at production scale: up to 24 agents sweep the finished storefront in parallel, each with its own brief: broken pages, wrong or mismatched data, and rare edge cases a sample check never reaches. The last mile to 100% quality.

The evaluation gate

Nothing reaches a live store without passing it

Fifty-four checks across nine categories, plus thirteen update-safety checks that run on any file touching products already live: schema integrity, image quality with a live pixel probe, text quality, content completeness, cross-link integrity, handle hygiene, numeric sanity, client rules, platform specifics, and update safety.

RULE 01

Only the changing columns go in the file. A column that is not in the file cannot be damaged.

RULE 02

Blank if there is no real source, never an invented value. A catalog that admits a gap is worth more than one that fabricates a plausible number.

WHAT IT PROTECTS

An update file can delete data it never even mentions. Thirteen checks stop that from happening, and no file reaches a live store without passing them.

The rule against invented data is enforced in code, not in a prompt.

02 / Built versus live

A queue of launches, not a backlog

Almost everything is built, reviewed and waiting on a client-side date: a store password, a domain cutover, a launch calendar. None of it is blocked on my side. Each row is a launch already queued.

Project
Built
Live today
What holds it
Jewelry
41,178 SKUs · 867 models · 1,586 pages
Final QA and designs done, behind a password
Client content team finishing descriptions. Earlier delays came from the ERP not being ready to receive a catalog of this shape, not from the build
Sports 2
48,438 SKUs · 1,715 products
Final QA and designs done, behind a password
Launch date not set
Sports 1
7,136 SKUs · 219 models · 701 products
156 products live
Launching in drops, drop one shipped. Next season already built and waiting on its launch date
Sports 3
1,148 products · 12,247 SKUs mapped
In production
Build in progress
Sports 4
3,253 products · 1,357 models
Initial mapping · recon
Queued behind brand 3
Watches
4,344 product pages · 12 luxury brands
Built and reviewed, in precision QA, behind a password
Awaiting launch · domain cutover
Designer boutique
6,788 product pages from 2,149 models
Final QA and designs done, behind a password
Client to run the import
03 / The engines behind the work

The machinery, with numbers attached

{{ shownCount }} of {{ totalCount }} systems
Source assessment · before any build starts · 2026 Recon: what exists, what is missing, what it will cost 534articles in the target list, live example 16content blocks scored 7page types defined 245gaps found before build ▾
Problem

Every project starts the same way, and nobody prices it: a brand name, source files in different shapes, a target list. What the data covers and what is missing is unknown. Quote before that is answered and you are guessing.

Solution

A recon pipeline reads every source at once and executes automated data reconciliation against the target schema. It tests compatibility across databases, measuring coverage block by block. This automated reconciliation is the primary factor determining project complexity and timeline: straightforward when data aligns, engineered from scratch when divergent (why #ProjectBanana took 6 months of modeling vs 2 weeks for mapped feeds). It joins stock report, linesheets, master data and asset library on a verified key, measures coverage block by block, sorts the catalog into page types, writes the build spec, and lists every gap with its reason.

Impact

On the live assessment: 534 articles and 2,384 SKUs scored across 16 content blocks, 7 page types, 40 fields specified, 14 mockups from real data. Coverage came back honest: barcodes 100%, color 98%, composition 69%, descriptions 50%, highlight icons 8%. The client master file covered 56 of 534, wrong season.

Full build pipeline · one adapter per brand · 2026 The build pipeline behind every catalog 17 dedicated pipelines 12 watch brands, each a model set 4 sports brands 1 jewelry catalog ▾
Problem

No single pipeline reads every brand. Twelve watch brands, four sports brands and a jewelry catalog each publish a different data shape, a different variant model and different bot defenses.

Solution

Not a scraper. One shared core with an adapter per brand: collection agents read the source, processors match models and expand variants in parallel, price and stock are enriched, and a writer agent assembles an operations-ready workbook. Checkpoint and resume let each brand run in waves.

Impact

The pipelines collected the source behind everything on this page: 4,344 watch pages, 41,178 jewelry SKUs and over 90,000 sports SKUs. A real run finds 3,502 products, matches 575 of 575 models and writes 11,082 variant rows in 17.3 minutes.

Watch a run →
Autonomous DAM ingestion · multi-server asset harvester · 2026 Autonomous DAM Ingestion: multi-server asset harvester 3 DAM architectures cracked 18,000+ master assets automated Zero API wait direct CDN & angle parsing ▾
Problem

Enterprise retail brands manage product media across closed, isolated DAM silos (Adobe Dynamic Media / Scene7, Bynder, and enterprise media hubs). Raw images, multi-angle variations (sole, toe, pair, details), and colorways lack unified export feeds, trapping merchants in days of manual downloading, mismatched variants, and broken links.

Solution

I built an autonomous extraction machine reverse-engineered for each server architecture. The engine deconstructs dynamic Scene7/AEM endpoint commands, probes and resolves angle parameters (DEFAULT, PAIR, TOE, SOLE, A), parses tokenized Bynder collections, applies perceptual hash deduplication (pHash), and structures master media into Shopify-ready schema.

Impact & Case Study

Over 18,280 high-res master assets harvested across 3,620+ articles with 100% color-to-angle accuracy. Completely eliminates manual image handling, reducing catalog asset ingest from weeks to minutes with zero broken links.

Parallel review operations · storefront sweeps · 2026 Multi-agent parallel review: finding the needle in the haystack 12-24 agents per sweep ~80% less processing time 14 watches recovered from one metafield error ▾
Problem

A finished storefront of tens of thousands of pages hides its own defects: broken images, wrong or mismatched data, gaps, and rare edge cases that no sample check will ever surface.

Solution

I run the same parallel pattern over a live storefront: 12 to 24 agents sweep the store at once, each with its own brief, hunting broken images, data that does not belong, structural gaps and edge cases, and reporting findings as structured rows.

Impact

The needle gets found in the haystack. On the watch catalog one metafield error was quietly keeping 14 watches out of the catalog altogether, and a code bug had left wrong attribute and price values on a handful of random pages. Neither shows up in a sample check. The sweep finds them, in a fraction of the time a linear review would take and with about 80% less processing time.

Feature, spec and filter layer · 2026 Feature mapping, rating scales and smart filters 5-point rating scales Per model spec extraction ▾
Problem

Raw catalog data does not make a store shoppable. Manufacturer feature language is inconsistent, specs live in prose, and filters built on free text either return everything or nothing.

Solution

Feature language normalized into a fixed set and matched to icon sets, technical specs extracted or written per model, comparable rating scales derived for cushioning, flexibility and support, and filter attributes generated from real structured values.

Impact

The same feature reads the same way on every product page, and a shopper can compare models on a measure rather than a paragraph.

Site architecture · 2026 Sitemap, category tree and internal linking Full site mapped per build ▾
Problem

A catalog with no architecture is a pile. Categories, collections, landing structure and links have to be designed or nothing is findable, by shoppers or by search engines.

Solution

Category tree and collection logic designed alongside the data, landing structure defined, internal linking generated, and old-to-new page mapping produced on replatforms so rankings survive.

Impact

Every build leaves with a navigable, indexable structure instead of a flat product list.

Certification gate · one per pipeline · 2026 The QA engine, tuned per pipeline 54 checks in the certification suite 9 categories, plus update safety 14 locked client rules Zero unverified fields shipped ▾
Problem

A catalog at scale is only useful if someone can prove it is right, and no two catalogs are wrong in the same way. Watches fail on terminology and placeholder prices, apparel on images bound to the wrong color.

Solution

A gate between build and handover, rebuilt per pipeline: images probed live for size and reachability, translation checked against a locked glossary, prices range-checked, variants and cross-links resolved, every editorial decision recorded as a rule. On the watch catalog: a 10-check mapper QC, 14 client rules, 54 certification checks, 13 update-safety checks.

Impact

Every brand went out on a full pass rather than a promise: 54 of 54 checks green, image dimensions verified on 4,028 of 4,028 images. Every catalog leaves with its own audit trail.

Taxonomy agent · 2026 Catalog classifier 98% classified automatically 2% flagged for second review ▾
Problem

Manual categorization of new products was slow, inconsistent, and did not scale across catalogs.

Solution

A fully automatic classifier that reads titles and descriptions, then maps products to categories, attributes, and taxonomy logic. Anything ambiguous or disputed is flagged rather than guessed, and routed to a second review pass.

Impact

98% of products classified with no human touch, the remaining 2% marked for a second review. A reusable pattern that carries to other catalogs.

Marketplace onboarding · 2026 Project Teapot: marketplace catalog automation 4 retailer templates 441 fields mapped 57 to 90 avg SEO score ▾
Problem

A homeware brand was expanding from one WooCommerce catalog onto 4 UK Mirakl marketplaces, each with different templates, taxonomy, image rules, and SEO standards.

Solution

An automation workflow mapping one catalog through Mirakl into 4 retailer-ready templates: normalization, category mapping, compliance fields, image structuring, retailer attributes, and SEO content.

Impact

4 marketplace-ready templates delivered, 441 fields mapped, 2,414 SEO text items rewritten, average SEO score up from about 57 to about 90.

Delivery layer · updates on demand · 2026 The publishing layer behind the sites ~95 live pages 0 release cycles ▾
Problem

A practice and a body of work both need a public face that stays current, without a redesign or an agency every time something ships.

Solution

An automated publishing flow with no version limit: pages go live on demand, plus a password-gated submissions site with per-page encryption.

Impact

Roughly 95 live pages to date across a brand site, an artist portfolio, and gated client material. New work appears as it happens.

Guarded edit agent · 2026 Safe-edit agent for a live site 26 pages migrated 66 sections mapped 563 assets ▾
Problem

A live site had to stay editable by a non-technical owner without any risk of publishing a broken page.

Solution

WordPress to static, with 66 sections mapped so any edit can be located by description, and an edit agent that clones to draft, snapshots, and publishes only on an explicit phrase.

Impact

Fast, cheap to host, and safe to edit. Live and maintained.

Lead tracking tool · 2026 Mini CRM application Live demo online ▾
Problem

Needed a lightweight pipeline-tracking tool, not a heavy Salesforce-style setup. Just status, notes, and next actions in one place.

Solution

A practical Mini CRM: lead status tracking, notes, next actions, and a simple follow-up structure.

Impact

A focused internal tool that keeps pipeline data organized, and a working demonstration of turning a personal ops problem into a usable application.

Open live demo →
05 / Services

Where I come in

Work starts with a system, not a task list. Design first, then the build: pipelines, connectors, the capability layer, and the checks around them. Every job opens with a recon that maps the source and sizes the work, then a fixed build plan. After launch the same machinery keeps the catalog current: new products, refreshed data, filled gaps. Project, monthly retainer, or full-time.

{{ s.num }}

{{ s.title }}

{{ s.body }}

{{ s.cta }}
06 / Stack

Tools actually in production use

Pipelines & automation

Python Multi-agent pipelines Checkpoint & resume Anti-bot scraping Apps Script Make.com Zapier

Commerce & marketplaces

Shopify Magento WooCommerce Mirakl Rithum / DSCO LogicBroker Odoo ERP Cymbio

Data & reporting

SQL Advanced Excel Power BI KPI dashboards Postman / APIs CSV · XML · JSON · EDI

How I show up

Owner of messy problems. Ambiguity in, something the next person can use out.

Team bridge. Sits between commercial, ops, product, data, and engineering: the role that translates.

Proof before handover. Nothing goes out on my word alone. Every build carries the checks that prove it and the record of how.

07 / Experience

Five years making other people's data behave

2026 to present
alfa · independent practice · remote

AI Solutions Architect, eCommerce catalog and data systems

Six client builds in parallel: 108,999 SKUs built and 15,134 product pages created, each build reaching a working pipeline and a first catalog in days, not months, then kept current with new products, refreshed data and filled gaps. Catalog engineering, pipelines that run many agents at once in production, translation held to a locked glossary, platform and ERP connections, Hebrew localization, feature and filter design, and the checks that make the output trustworthy.

2025
Marcella New York · contractor

eCommerce & Marketplace Manager

Owned marketplace data flows across Shopify and third-party retail channels. Inventory buffers and safety-stock logic to prevent overselling, taxonomy alignment, pricing, retailer compliance, and content work.

2022 to 2025
Cymbio

Data Implementation Manager

Mapped complex brand catalogs to strict retailer schemas (Mirakl, Nordstrom, Macy's) across EDI, CSV, API, XML, and JSON. SQL and Apps Script tooling cut manual operational work by about 60%. SOPs and intake forms improved partner onboarding by about 50%.

08 / Get in touch

Bring the problem. I will design the system.

A platform that will not connect, a catalog that does not exist yet, a catalog going stale, or a workflow still running on people. Open to a project, a monthly retainer, or a full-time role. Reply within one business day.

Goes straight to my inbox. Reply within one business day.