STEP earns its place in a product domain for a reason that is easy to state and easy to underrate: it thinks in hierarchies. Most platforms in the category model a product as a record with attributes and treat the tree as navigation. STEP treats the tree as the model. Families, products, variants and SKUs are levels with their own attribution, classifications are first-class structures alongside the product hierarchy rather than tags upon it, and language and market variance is handled by a dimension mechanism designed in from the start rather than by cloning fields. A product catalog is never finished, and the things that change most often are the structure, the assortment, and the channels it has to feed. A platform whose primary object is the structure suits that unusually well.

Product data still arrives the way product data always arrives: supplier spreadsheets in eleven incompatible templates, a GTIN column that is right most of the time, attribute names that mean different things to different category managers, images named by whoever exported them, and a channel deadline that does not move because your data was not ready. Turning that into channel-ready records is a design exercise, and this paper is the design. It assumes STEP on the current SaaS line, a real catalog with live supplier feeds and more than one downstream channel, and a stakeholder who will eventually ask why time-to-shelf has not improved. It names the decisions that are expensive to revisit, and it is specific about where the reference architecture reaches beyond the platform, because those integration points are where effort estimates usually go wrong.

One piece of context before anything else. Stibo Systems enters this design as a stable counterparty: headquartered in Aarhus, foundation-owned through Stibo Software Group, which gives the roadmap a longer horizon than a private-equity clock, and named a Leader in the 2026 Gartner Magic Quadrant for Master Data Management Solutions. The platform now ships on a rolling update cadence, with the 2026.3 update landing in September, and the company has spent 2024 through 2026 layering AI into the workflows rather than beside them. All of that is welcome. None of it changes the discipline: design against what your update actually ships, and treat the roadmap as upside rather than a dependency.

1. What shipped between 2024 and 2026, and what to do about it

The recent release history matters to a blueprint because several items on it change where effort goes. The 2024.3 update brought AI-assisted mapping for syndication, which cuts the most tedious task in a PDX channel build from days to hours. The 2024.4 update added an AI-powered address matcher, a pre-trained model for customer and supplier address data that reduces the manual review load on party matching. The July 2025 rollout introduced the Stibo Intelligent Assistant, which generates SEO-oriented product descriptions and translates them across markets at scale, alongside machine-learning matching in the Customer Experience Data Cloud built on pre-trained networks that learn from steward feedback. And in June 2026 Stibo launched its MCP Server, exposing the semantic data graph, its relationships and its hierarchies to AI agents through a governed interface, with a Microsoft Copilot Studio and Fabric integration demonstrated at NRF in January showing what a first-party shopping agent grounded in mastered product data looks like.

Read as an architect, the list sorts into three piles. The AI mapping and the address matcher are build-phase accelerators: plan to use them, and reclaim the estimate lines they replace. The Intelligent Assistant is an operating-model question, because generated product copy needs governance before it needs enthusiasm, and section 7 gives it the treatment it deserves. The MCP Server is an interface you will be asked about within a year of go-live, whether or not you plan for it, so section 11 plans for it. What none of them are is a reason to change the foundations, because the foundations of a STEP implementation are the hierarchy, the dimensions and the attribution model, and those are decided by your business, not by a release note.

Table 1. The 2024 to 2026 capability drops, sorted by what they change in your plan.
CapabilityWhat it isWhat to do about it
AI-assisted syndication mappingProposes channel mappings in PDX instead of a human building them field by fieldUse it, then review every proposed mapping. It accelerates the work; it does not own it
AI address matcherPre-trained model for matching customer and supplier addressesAdopt for the supplier and party tables. Validate against a labelled set before trusting it with merges
Stibo Intelligent AssistantGenerates SEO-oriented product descriptions and translations at scaleGovern before you generate: provenance flags, review workflow, and a hard boundary around regulatory fields
ML matching, CX Data CloudPre-trained networks with rule training and learning from steward feedbackRelevant if the same programme masters customers. Measure precision the same way as any matcher
MCP ServerGoverned agent access to master data through the semantic graph, launched June 2026Treat as a published interface: read-only to start, scoped permissions, and a named owner
Rolling 2026.x cadenceQuarterly-style updates rather than monolithic versions, 2026.3 arriving SeptemberPin your design documents to the update you are actually on, and re-verify AI feature availability per update

2. What a STEP repository actually is

A handful of words carry most of the design weight in STEP, and getting them straight in the first week saves an argument in the third month. Everything is a node in a hierarchy. The product hierarchy is the primary structure: it owns the products, it defines the levels, and attribution validity hangs from it, so an attribute is declared valid for the nodes where it applies rather than existing globally. Classifications are separate hierarchies used to organise the same products for other purposes, merchandising, procurement, web navigation, and a product links into them by reference rather than living in them. Attributes come with types, validation and units; lists of values are governed objects in their own right rather than free text with good intentions; and reference types link products to assets, to classifications and to each other. Assets, images, documents and rich content, are managed objects with their own hierarchy, not attachments.

The single most expensive modelling decision in a STEP programme is where the golden record lives, which in a hierarchy-first platform means which level is the unit of truth. A family carries shared marketing copy, a product carries the attributes common to its variants, a variant carries colour and size, and a SKU carries the logistics. Put channel-facing content too low and you maintain the same description two hundred times; put logistics too high and every variant claims the same weight. The discipline is to write, per attribute, the level at which it is asserted and the levels at which it is inherited or derived, and to defend that document in a workshop with the merchandising and supply chain owners before any data is loaded. Restructuring a populated hierarchy is possible in STEP, and it is also the closest thing the platform has to open-heart surgery. It is far cheaper to argue about levels on a whiteboard.

Two honest notes belong in the same breath, because peer reviews of the platform repeat them with remarkable consistency. STEP is powerful and stable, the match engine is well regarded, and the modular design rewards teams who model deliberately, and the initial setup is genuinely complex, with a learning curve that punishes an under-skilled build team. Budget for people who have done it before. And the analytics and self-service reporting inside the platform are not where the product shines, so if your stakeholders expect dashboards over the catalog, plan a reporting export to the BI estate as a workstream from the start rather than promising it from the workbench. Neither point argues against the platform. Both argue against staffing it thin.

3. Dimensions and contexts: the decision you cannot cheaply reverse

STEP handles variance across languages and markets through dimensions, and the mechanism repays being understood precisely because the vocabulary misleads people. Data is stored in dimension points, not in contexts. A context is a filter, a combination of exactly one dimension point from each dimension, language and country being the usual pair, and selecting a context in the interface changes which data you see, not where data lives. To make anything vary, the object must be declared dimension dependent: object names for translated product titles, attributes and their lists of values for market-specific content, and reference or link types where a product should carry a different image or document in a different market. A reference type can depend on a single dimension point only, so if you need variance across more than one dimension, that is a new reference type, not a setting.

The vendor's own recommended practice says the important thing plainly: think hard about dimension dependency when setting up the system and before data is migrated in, because it is easier to remove a dependency than to add one after references exist. Take that seriously. The corollary guidance is just as useful in the other direction: if only a handful of attributes vary, for products sold in only a few countries, separate attributes such as a description per language are the simpler and cheaper design, and dimensions earn their complexity only when variance is broad. One requirement forces your hand regardless of preference: the Stibo GDSN solution requires each target market to reference a context containing a language and a country dimension, so a programme with GDSN publication in scope designs its dimension set around the target market list on day one.

Table 2. Dimension dependency, decided object by object before migration rather than discovered after it.
ObjectMake it dimension dependent whenCaution
Object namesProduct titles are translated, or differ by market for legal or brand reasonsA missing name in a context displays as the raw ID, so completeness checks must cover every live context
Attributes and LOVsValues genuinely differ across many markets: descriptions, compliance text, market-specific claimsFor a handful of varying attributes in a few countries, separate per-language attributes are the simpler design
Reference and link typesA product needs a different asset or document per market, packaging shots being the classic caseOne dimension point per reference type. Decide before references are created, since removal is easier than addition
ContextsOne per real language and market combination you serve, and no moreEvery context multiplies completeness work and export testing. Add them when a market is real, not aspirational
GDSN target marketsAlways, if GDSN is in scope: each target market must reference a language-and-country contextThis is the one place the dimension design is dictated to you. Start from the target market list

4. Model attribution so completeness is computable

The purpose of a product master is to answer one question per channel per market: is this product ready. Readiness is a computation over attribution, which means the attribution model has to be designed for computation rather than for tidiness. Three rules do most of the work. First, bind attribute validity to the hierarchy and the classifications deliberately, so that a completeness check on a power tool never asks about thread count and a check on bed linen never asks about voltage; validity is your first line of data quality because an attribute that cannot exist can never be wrong. Second, prefer governed lists of values over free text everywhere a downstream channel will filter or facet on the field, because a colour facet built over free text is a merchandising incident scheduled in advance. Third, make units explicit and single, one unit per attribute with conversion at the boundary, because a weight field holding grams from one supplier and kilograms from another passes every format validation and fails only in the customer's basket.

Then define readiness itself as data, not as folklore: a named set of required attributes, references and assets per channel per market, evaluated continuously, surfaced to the people who can fix the gaps, and used as the gate in the syndication flow so that incomplete products are held rather than published thin. Business rules are the enforcement mechanism, conditions to evaluate, actions to normalise and derive, and they belong at the point of write so that non-conforming data is rejected or routed rather than reported on a month later. The test of a good attribution model is that a category manager can be told, in one screen, exactly which three fields stand between a product and its launch on a named channel. If your model cannot produce that sentence, it is not finished, however elegant it looks.

5. Onboarding: suppliers first, spreadsheets second

Inbound product data has three doors, and the design job is to route each source through the right one. Supplier-provided content belongs in a PDX onboarding channel, where suppliers deliver against your data requirements rather than their export habits. Bulk and system-to-system loads run through Import Manager and inbound integration endpoints, where mappings are declared, business rules clean and normalise on the way in, and an IIEP with the match-and-merge importer handles party-shaped data, supplier and manufacturer records, with thresholds that merge confident matches, create unique golden records, and route the uncertain middle to a clerical review workflow. And GDSN subscription handles the trading partners who publish through the pool. What should not exist, after the first quarter, is the fourth door: the spreadsheet emailed to a steward who keys it in, because every record that enters that way bypasses every rule you have written.

The PDX onboarding channel has configuration details that decide whether multilingual supplier content arrives cleanly, and they are worth knowing before the workshop rather than after it. The Context ID parameter on the channel is mandatory and sets the default market and language for extraction and load; structural properties such as hierarchy names, attribute names and descriptions, validations and list values do not vary by language in PDX and are drawn from that default context's language. Suppliers onboarding multiple languages into one market use the language handling and language mapping attributes; suppliers onboarding into multiple markets need the market dimension parameter and supplier contexts configured per supplier classification. None of this is difficult. All of it is the kind of detail that, discovered in build, turns a two-week channel setup into a six-week one.

The sequencing rule from every serious MDM programme applies here unchanged: load everything, reconcile the counts, and only then match and merge. Matching supplier or manufacturer records against a partially loaded table produces clusters built on absent evidence, and those clusters leave behind merge decisions that have to be purged and redone. Treat the reconciliation report, row counts and null rates against source extracts, as a signed deliverable, and keep the load-and-match schedule serialised rather than letting per-source jobs race each other.

Table 3. Inbound paths into a STEP product hub, and what each one is for.
PathUse it forDesign notes
PDX onboarding channelSupplier-provided product content, delivered against your requirementsContext ID is mandatory and sets the structural language. Plan language and market handling per supplier segment
Import Manager and IIEPsBulk loads, ERP and system feeds, migrationsBusiness rules normalise on the way in. Chunk large loads and treat batch size as a tested number, not a default
IIEP with match-and-merge importerParty-shaped data: suppliers, manufacturers, and customers if in scopeAbove threshold merges, unique records golden, the middle goes to clerical review. Size the review queue honestly
GDSN subscriptionTrading partners who publish through the data poolTarget markets dictate contexts. Validate pool data with the same rules as any other source
Emailed spreadsheetsNothing, after the first quarterEvery record entering this way bypasses every rule you wrote. Route the sender to a channel instead

6. Identity: mostly deterministic, and honest about the rest

Product identity is kinder than customer identity, and the design should exploit that rather than importing a probabilistic apparatus it does not need. A GTIN, where present and sane, is an exact key. Manufacturer plus manufacturer part number, normalised, is the next tier. The design work is not in the comparison, it is in the normalisation before the comparison, and business actions are where it lives: strip and standardise part-number punctuation, canonicalise manufacturer names against a governed list, and refuse placeholder values, the row of nines, the single x, the word unknown, before they enter any comparison, because a placeholder shared by four hundred records is how a four-hundred-record cluster is born. Where deterministic keys run out, at the long tail of supplier-created records with no GTIN and creative part numbers, the fuzzy tooling earns its keep, and the same doctrine we apply on every platform applies here: build a labelled set from your own data, measure precision and recall against it, treat a false merge in the labelled set as a release blocker, and never switch on automatic merging that has not earned its precision in writing.

Survivorship deserves the same explicitness. Golden record survivorship rules in STEP let you configure, per attribute, which source wins, and the ranking should be built field by field with the data owner: the ERP wins the logistics attributes it is the system of record for, the supplier wins the technical specification they manufacture to, the content team wins the copy they write, and a steward's correction outranks an inbound reload except on the short list of fields that must mirror an external registry. Write the ranking down as a document, because it becomes the most consulted artefact of year two, and reconstructing it from configuration screens eighteen months later is miserable work. For supplier and customer party records, the AI address matcher introduced in 2024.4 is a genuine reviewer-workload reduction, and the machine-learning matching in the Customer Experience Data Cloud is worth evaluating if the same programme masters customers, with the caveat that a learned matcher is still a matcher: it gets a labelled set and a scorecard, or it does not get production.

7. The Intelligent Assistant writes fast. Decide who approves.

The Stibo Intelligent Assistant generates branded, SEO-oriented product descriptions and translates them across markets at scale, and on a catalog with thousands of thin records it is the single largest productivity lever in the platform. It is also generative, which means the governance has to exist before the volume does. Three rules make it safe. First, provenance: every generated or machine-translated value carries a flag saying so, as an attribute, so that a channel, an auditor or a lawyer can distinguish asserted content from generated content two years from now. Second, review: generated copy enters the same approval workflow as human copy, gated by the same readiness rules, and nothing generated flows to a channel unreviewed until a category has demonstrated, with a sampled error rate, that it has earned lighter-touch review. Third, a hard boundary: regulatory and safety content, ingredients, hazard statements, compliance claims, certified specifications, is never generated and never machine-translated without qualified human sign-off, full stop. Position the assistant as a drafting colleague for the content team rather than a replacement for it, and it takes real friction out of enrichment. Position it as a way to skip the content team, and you will syndicate your first hallucinated specification within a quarter.

The boundary with specialist services stays where it has always been. Postal and address verification for supplier and customer records, translation memory for regulated markets, and image quality services remain the business of providers who do only that; STEP masters, governs and orchestrates. The assistant standardising inconsistent casing and drafting a bullet list is the tool used well. The assistant validating an address or asserting a compliance claim is the tool used wrongly, and the design document should say so in those words.

8. Workflows: new product introduction is a production line

STEP's workflow engine is built for exactly this domain: states, tasks assigned to roles, deadlines, and events that move products through introduction, enrichment, approval and publication. The design mistakes are the same everywhere. The first is building a workflow per scenario until thirty models exist and nobody dares change any of them; four or five models, new product introduction, enrichment and translation, change approval, duplicate and exception handling, periodic review, carry a product domain for years, with the variation living in conditions rather than in separate models. The second is designing screens instead of capacity. A workflow is a queue, and a queue has arithmetic: expected products per week, times tasks per product, times realistic minutes per task, compared honestly against the people you have. If the arithmetic does not close, the fix is prevention, better supplier requirements in PDX, stricter validation at the door, generated first drafts from the assistant, not exhortation. Instrument time-in-state from the first day, because the state where products sit longest is the constraint, and the constraint is where the next quarter's improvement budget belongs.

9. Syndication: OIEPs that survive peak season

Outbound is where a product hub is judged, and STEP's outbound integration endpoints reward the same discipline the vendor's own performance guidance prescribes. Give each significant consumer its own dedicated OIEP rather than sharing one, and give important integrations separate queues, so that the commerce feed is never waiting behind a bulk extract. Event-based endpoints publish changes as they happen, with business conditions as event filters so consumers receive the changes they care about rather than everything; scheduled endpoints serve the consumers who want a nightly world. Batch size is a tested number, not a default; cross-context exports beat exporting per context serially; and the guidance to limit additional data and the number of templates per endpoint exists because every violation of it is invisible at test volume and expensive in November. If the in-memory component is in your licence, exports are one of the places it earns its keep, and peak-season export duration is a number to measure in performance testing rather than discover in production.

For PDX syndication specifically, one migration item belongs on the plan and not in the backlog: the current integration approach is the STEPXML-based OIEP, which replaces the older JSON-based integration and does not rely on locale-based mappings. Moving to it creates new context-based language layers in PDX which must be re-mapped on the channels, the sensible cutover runs both integrations in parallel against PDX pre-production and compares results before switching, and PDX production should only ever be the target of STEP production. It is a contained piece of work with a documented path, and doing it deliberately, in a quiet quarter, is how you avoid doing it accidentally in a busy one.

Table 4. Outbound practices, drawn from the platform's own performance guidance and from Novembers we would rather not repeat.
PracticeWhy
Dedicated OIEP per significant consumerIsolation. One consumer's bad day stops being every consumer's bad day
Separate queues for important integrationsThe commerce feed should never queue behind a bulk extract
Event filters via business conditionsConsumers receive the changes they care about, and event volume stays proportional to change
Tested batch sizes, multithreading where it fitsDefaults behave at fifty thousand records and misbehave at five million
Cross-context exports, limited additional data and templatesPer-context serial exports and template sprawl are the classic silent multipliers of export duration
STEPXML-based OIEP for PDXThe current integration path. Migrate off the JSON-based approach in parallel against pre-production, then re-map channel language layers
Readiness gate before publicationIncomplete products are held, not published thin. The gate is the completeness computation from section 4

10. Agent access is an interface you publish, not a feature you enable

The Stibo MCP Server, launched in June 2026, gives AI agents governed access to master data through the semantic graph, its relationships and its hierarchies, and the Microsoft integration shown at NRF makes the direction concrete: shopping and service agents grounded in mastered product data rather than in whatever a crawler found. The pitch is sound, agents are only as good as the data they act on, and the design posture should be the same one you would take for any new consumer of the hub, because that is what an agent is. Start read-only. Scope the exposure to the published, channel-ready projection rather than the whole repository, since a work-in-progress record that a steward has not approved is not something an agent should be answering questions from. Give the interface a named owner, a versioning story and a deprecation policy, exactly as you would a REST contract. And log agent queries from day one, because the questions agents ask are a free and unusually honest map of which attributes actually matter to consumption, which is information your attribution roadmap can use. Treat it this way and the MCP Server is the cheapest new consumer you will ever onboard. Treat it as a switch to flip and it is a governance incident with a press release.

11. Sequencing, and how to spend the first ninety days

The first fortnight is for the decisions that are expensive to revisit, not for modelling. Pin the update you are building on and re-verify which AI capabilities your tenant actually has. Write the hierarchy-level document: which level asserts which attributes, defended in a workshop with merchandising and supply chain in the room. Decide the dimension set from the real market list, GDSN target markets first if they are in scope, and put in writing which objects are dimension dependent, because that decision hardens the moment data is migrated. And start the supplier segmentation for PDX onboarding, since the suppliers' calendars, not yours, are the long pole in every onboarding plan.

Weeks two to six belong to the model and to one honest load. Attribution validity bound to the hierarchy and classifications, lists of values governed, units single and explicit, and the readiness definition per channel per market written as data. In parallel, load one real source end to end, ERP or the largest supplier, reconcile the counts, and sign the reconciliation. One source loaded and counted teaches you more about your data than four weeks of workshops, and it de-risks the ingestion design while there is still time to change it.

Weeks six to twelve are where the programme earns or loses its credibility: the identity and survivorship work, the first PDX supplier channel live with a friendly supplier, the first OIEP publishing to a non-production consumer against the readiness gate, and the workflow arithmetic computed from the observed suspect and exception rates rather than from hope. If the steward capacity number and the headcount do not meet, say so in that month, because the levers, thresholds, prevention, supplier requirements, are all still cheap to move. And go live with the same posture we argue for on every platform: automatic merging off, generated content fully reviewed, and both earned back rule by rule and category by category as the evidence arrives. It costs a few weeks of extra effort. It buys the ability to say, with evidence rather than assertion, that the hub has never published a record it should not have, and that sentence is worth more to the programme than any feature you could ship in those weeks.

12. Design checklist

These are the decisions we would take to a steering committee on a STEP product programme, in roughly the order they need answering. Most of them are not platform questions at all. They are choices about how the reference architecture fits your organisation, and every one of them is cheaper to make in design than to discover in test.

Table 5. The design checklist for a STEP product MDM programme.
DecisionWhy it mattersHow we handle it
Golden record levelRestructuring a populated hierarchy is the platform's open-heart surgeryA per-attribute assertion-level document, defended in a workshop before load
Dimension dependencyEasier to remove than to add once references exist, and GDSN dictates part of itDecided object by object before migration, starting from the real market list
Readiness as dataTime-to-shelf is decided by whether completeness is computable per channel per marketRequired attributes, references and assets per channel, evaluated continuously and used as the publication gate
Supplier onboardingSupplier calendars are the long pole, and PDX context configuration decides multilingual cleanlinessSegment suppliers in week one, first friendly channel live by week ten
Identity and survivorshipDeterministic where keys exist, measured where they do not, and written down either wayNormalisation via business actions, a labelled set for the fuzzy tail, a field-by-field trust ranking document
Generated content governanceThe Intelligent Assistant scales output; only governance scales trustProvenance flags, review workflow, and a hard boundary around regulatory fields
Syndication architectureOutbound is where the hub is judged, usually in NovemberDedicated OIEPs, separate queues, tested batch sizes, STEPXML-based PDX integration, readiness gate on every channel
Reporting expectationsSelf-service analytics inside the platform is a known weak spotA scheduled export to the BI estate, scoped as a workstream from day one
Agent accessThe MCP Server will be asked about within a year, planned for or notRead-only, scoped to the channel-ready projection, named owner, logged from day one

13. The eight things we would insist on

First, decide the golden record's level in the hierarchy before any data is loaded, per attribute and in writing. It is the one decision in a STEP programme with no cheap undo, and every other design choice inherits from it.

Second, settle dimension dependency object by object before migration, because it is easier to remove a dependency than to add one after references exist, and use dimensions only where variance is broad. A handful of varying attributes in a few markets wants separate attributes, not a dimension.

Third, define channel readiness as data, evaluated continuously and enforced as the publication gate. The test is one sentence: a category manager can see exactly which fields stand between a product and its launch on a named channel.

Fourth, close the spreadsheet door. Suppliers deliver through PDX channels against your requirements, systems deliver through endpoints with business rules at the boundary, and by the end of the first quarter there is no fourth path.

Fifth, keep product identity deterministic where keys allow and measured where they do not. Normalise before comparing, refuse placeholders at the door, and let no matcher, classical or learned, merge automatically until its precision has been demonstrated against a labelled set built from your own data.

Sixth, govern the Intelligent Assistant before scaling it: provenance flags on generated values, generated copy through the same approval workflow as human copy, and regulatory content permanently out of bounds. A drafting colleague, not a bypass.

Seventh, build syndication for peak season on day one: a dedicated OIEP per significant consumer, separate queues for the integrations that matter, tested batch sizes, the STEPXML-based PDX integration rather than the legacy path, and export durations measured at real volume before November measures them for you.

Eighth, treat agent access as a published interface. Read-only to start, scoped to the channel-ready projection, owned, versioned and logged. The agents are coming to your product data either way; the only question is whether they arrive through a contract or through a gap.

STEP rewards teams who lean into what it is: a hierarchy-first, model-driven platform with a strong match engine, a mature supplier and syndication ecosystem, and, as of this cycle, a credible AI layer inside the workflows rather than beside them. Play to that and you get a product hub whose structure the business can actually change, whose completeness is a number rather than an opinion, and whose channels receive products the moment they are ready rather than the moment someone notices. The work that surrounds it, the level decision, the dimension design, the readiness definition, the labelled set, is the same work every serious product MDM programme does on any platform, and all of it is far cheaper to plan than to retrofit. If you are standing up a product domain on STEP, or you have one that is technically live and not yet fully trusted, we are happy to talk it through.

Found this useful? Share it.
LinkedInX / TwitterEmail