Who owns the product data | Ep. 11

Watch the podcast now

Episode 11 – Who Owns the Product Data? – Transcript

Matt: Welcome to Data Versus Commerce, where we explore the messy middle between database and doorstep. I’m Matt Johnson.

Floyd: And I’m Floyd Blaikie. Let’s dig in. Okay, the last time we had Jay on the podcast, we talked about hitting a ceiling — in this context, that’s the moment you realize as an e-commerce leader that you’ve kind of outgrown your tools, and your whole operation suddenly seems to be held together with workarounds. We ended that episode on a question we didn’t actually answer — in the business, we call that a hook — is it the platform that’s holding you back, or is it the data underneath it? So today we’re going underneath, because your product data has turned into the thing that decides whether an AI recommends you, whether a customer can ever find a replacement part, whether you show up in that consideration set at all. A lot of teams out there have no real idea what’s going on down there in the foundation. So if you caught part one, you already know our first guest, Jay Roxe, who runs marketing at inriver and who is, as far as I’m concerned, just one exclamation point short of a superhero name. Jay, welcome back to the show.

Jay: Floyd, great to be back.

Floyd: And new to the show today, but I’m very glad to have him — Willem, Pivotree’s director of data platforms and services. Jay’s the guy who thinks about the PIM, Willem’s the guy who thinks about everything holding it up — the integrations, the single source of truth, all the stuff you don’t notice until it breaks. Willem, thanks for being here too.

Willem: Yeah, it’s good to be here, good to talk with you both, Jay and Floyd.

Floyd: All right, quick one before we get into it — what’s a product data horror story you’ve seen or been part of that still keeps you up at night?

Willem: I know of a company — I won’t name it — that ended up with a disgruntled employee who, before they left, pushed a whole bunch of terrible content through a solution, and it ended up online. People were scrambling to take it down. That was an unfortunate event.

Jay: One of the things that was the matter was they had no controls over who input data, who reviewed it, who signed it off, who approved the publication — a good example of data governance gone awry. I’ll leave that one in the horror-story category, and I can see why it keeps you up at night. I’ll go in a different direction — I know of a company that had a $250 million hole in their revenue line because they couldn’t actually publish the data on their replacement parts. It was just unavailable, so everybody who needed to fix these rather expensive machines was going out and getting the information, or the parts, through the gray market. That one’s less “I need governance to control rogue employees” and more “the foundation was just never solidified to begin with.”

Willem: Yeah, and product data management is hard enough as it is — but with spare parts, you also have to understand the relationships between the part and the thing it fits on, the fitment information, whether there are interchanges that can substitute for a part. It becomes much more complicated.

Jay: And where it came from to begin with, because supply chain tracking is such a thing now. We have parts we track for customers where we need to contain the product information that, say, for a certain number of years it was made in the United States and then got moved to being made in a different country — that didn’t change what the part looks like or where it fits, but it may change some of the downstream ramifications.

Floyd: A lot of layers here — and Jay, you called this a layered project, hitting the PIM ceiling, or if you’re not on PIM, hitting the Excel ceiling, in the last part of this series. So let’s take some of those layers apart. Willem, when you say “data foundation” in reference to PIM or any commerce platform, what are the actual layers, and which one do teams most often skip or under-resource?

Willem: Yeah, so there’s the information supply chain, as I call it — where does data come from, where does it need to go, who are all the parties involved, and how do you control what they do? If you’re a manufacturer, you need to control the product management team, the engineering team — it routes through marketing, out to e-commerce, print, supporting systems for e-commerce, distributors, data aggregators, and you need it for AI. So there are a lot of channels that need data in different ways. And on the upstream side, you have people who all think they know best, they do it their own way, and you need to herd that bunch of different personalities and different groups. That’s a huge part of the problem. If you’re a retailer or distributor, you’re dealing not just with your own organization, but with lots of companies you don’t really have a lot of control over, even though you’d like that control. They all think about their products their own specific way, even though you need to normalize it and present it cohesively on your e-commerce site. Those are some of the challenges you’re dealing with — and that’s not even talking about the systems that can help. That’s just the core business challenge of people doing their thing.

Jay: We think of it as product information orchestration, because you’re not just managing it — you have to get it from point A, enrich it, figure out what it needs to be for each channel, figure out where you’re sending it, who’s doing each action, and then get it out. When we talk with some of the large manufacturers, they almost wish they had as much control over some of their own upstream organizations within the same company as they would over a supplier to a distributor. The challenges continue to multiply as things get larger.

Floyd: It sounds like there’s a battle shaping up over who’s in charge of the product data. In your perspective, either one of you — should it be one person, one layer, or does the stewardship change as you move through the different parts of that information supply chain?

Jay: We did a survey recently that showed, when you start talking to organizations, nobody agrees on who should own the product data. Ask a head of IT, they think they should own it. Ask a head of marketing, they think they should own it. At a mature organization, they may have actually figured out the relationship. Floyd, I think there are actually two questions in what you’re asking — there’s who owns it, meaning who’s ultimately accountable for what comes in and goes out, and that’s all over the place. And then there’s who is responsible or allowed to touch it at each phase of its journey, and figuring out all the places it should go out. Both of those are things most organizations struggle with, and it’ll be multiplied as they move into an era where AI, rapid publication cycles, and the need to drive AEO is changing the information, and how it goes out, at a faster speed than many organizations have dealt with previously.

Floyd: So it sounds like there’s still a debate on who owns the product data. But Willem, how does a team actually know whether they can trust their product data? Even if the ownership question is fully sorted, and everyone knows their role — what are the red flags that tell a team, “I don’t think I can trust this, this might not be correct”?

Willem: Yeah, we do those audits for a lot of companies, and it starts with taxonomy — how do you define, group, and categorize your products? Every product has a different type. But the quality is discoverable in a number of different ways — there are dashboards that let you look at fill rates, meaning this attribute is filled out, this one is lacking values. There are dashboards that show outliers. There are ways to use AI — if data comes in from a supplier pool, you can do anomaly detection on individual attributes, but also attributes that may correlate to another attribute of that same product, and decide, well, if the description says this, then this attribute can’t possibly be value X or value Y. So there are ways, inside the workflow, to determine quality and immediately bounce a record back to the supplier, or fix it through an automated process — or you do it through dashboarding, and present bulk sets of data to an administrator or data governance person experienced at fixing those issues.

Floyd: Jay, you probably see a lot of the before and after — I imagine you see how transformational it is when a team has that kind of reporting, that dashboard where they can see if the data’s a little rotten or looking good. But what about the people still running on Excel? Excel’s not coming back with Clippy saying, “Hey, it looks like your data is bad, do you want help with this?” Practically, what are the symptoms when they don’t have that reporting capability, and they want to know, “Is my data working for me or against me?”

Jay: I think a lot of organizations are at the point where they’ve developed so many workarounds, so many places where they have one person who knows how the system is supposed to work, and they’re quietly praying that person doesn’t win the lottery and move to the Canary Islands. But I think what organizations start to see is they’re not as ready to put information out there. When we talked to about 400 manufacturers and distributors, we saw that only about a third of them have more than 85% of their publish-ready SKUs available at any given time. That’s not necessarily correlated just to Excel, but certainly Excel makes it a lot harder to move toward a solid data foundation, because there aren’t the tools — you’re not seeing Clippy, or these days Claude, or an AI tool built into your environment that creates underwriteable, trademarkable, brand-ready content updating anything that’s in Excel. People can go through and change it, but that’s by nature a slower, more manually intensive process, prone to errors.

Willem: Yeah, maybe we need Clippy 2.0 — I don’t know what that would be, but —

Floyd: Oh my gosh, I was just thinking, wow, I’m so grateful they didn’t turn Clippy into an AI bot, because I think I would log off forever if that happened. No, thank you.

Jay: I know what I’m sending you for Halloween.

Willem: You could vibe code it, Jay, and then send it to Floyd for her birthday — that’s a great idea. But one of the things we’re working on is plugging AI in at the edge of PIM. There’s a real change going on where PIM isn’t just a tool inside your organization — together with an MCP server, PIM can become something that governs your entire ecosystem. You can get all your distributors, as a manufacturer, to come in and pick up content with agents. Most people aren’t there yet, but it’s going to change very quickly — having your distributors pick up the data without you even needing to touch or change it. You can have an agent come in and do the consuming it needs to do. There’s no need for third-party data aggregators that morph that data into the target channel’s repository taxonomy — I know that’s a lot of words, but basically, the data doesn’t need to change for consumption by somebody else. Those folks just come shop for the data with you, and I think that’ll make huge changes. It also means you need to get your data in tip-top shape — you can filter down and only give people access to certain portions, but that portion needs to be perfect. You’ll be thinking less and less about readying data for publication — a slice of it will just be ready for the pull from external parties.

Jay: There’s a part that’s important for people to always remember: the degree to which that data will continue to change over time — not just the number of attributes, but the number of pieces you need to maintain around that data. Everything looks clean on launch day, and then you get a new channel that requires different information, or you start sourcing it from a different place — keeping up with that degree of change is a core part of making sure the foundation is solid. But I wanted to touch on something you said first — we were joking about what I’m now going to vibe code for Floyd, and Floyd, you may want to be aware, be prepared for that.

Floyd: I’m ready.

Jay: We’ve always focused on the idea that a PIM needed to provide a very extensible infrastructure, and I’ll say, even on top of that, releasing an MCP server and letting people code against it and build their own orchestration, their own agents, really creates a very strong system of work around that product data — not just a system of record.

Floyd: Yeah, that was your big idea last time, Jay — that the PIM is going to evolve from a place where data lives to a place where work gets done. So I’m hearing maybe two different things here — help me untangle these threads. What I hear from a lot of people in the space is if you point AI or an agent at bad data, you’re just going to get worse stuff faster, which I think is true. But we’re also talking about applying AI to rapidly changing data to stop it from getting bad. So is there a place at which your data is good enough to start pointing AI at it? And how do you know if you’re there?

Willem: Well, there are a lot of places where AI is starting to get used — in the creation of a PIM or MDM solution, in needs gathering, in development, in delivery, in blueprinting — all of that is now, at least for us, sort of AI-driven. And even in implementation, and support, triaging issues that come in. That’s where AI is being used. And then there’s a whole slew of little things you can do inside a PIM or MDM solution — we talked about anomaly detection from data coming in from suppliers, translation, keyword insertion, copy generation, extracting content from labels, data enrichment, as well as recognizing errors and fixing them, which is what you were bringing up, Floyd. Now, I would never — at least not yet — tell people to just let an agent or an AI process loose without some human-in-the-loop controls. But that is happening today. You can definitely set that up — if you can recognize an issue, it’s also easy to create a process where, in most cases, you can fix the things being noticed.

Jay: And I agree that human-in-the-loop validation is a key piece. But Floyd, I want to go back to your question of when you can start looking at it. As with any data process, understanding the data, figuring out how you want to model it, making sure the data fits into that model, adjusting, cleaning, normalizing — all of these are things that need to happen with both people and machines in the form of AI. It can accelerate it, but I’ll totally agree with Willem — we’re not yet at the point where you can say, “Okay, just go.” The other piece, though, is that as you work with people who’ve done more and more implementations, that pattern recognition — and the way it gets built into the tools or the process — gets faster and better. So there’s a learning curve that rapidly accelerates what can be done, by whom, and how fast.

Floyd: So the obvious next question, I think, is how do you actually do it — which I think has to be part three, because I can’t stop inviting Jay back on the show, and we could probably stay here for another hour, but we’ve got to land it at some point. So it’s time for the free pluggables. Jay, for the folks who want to keep going on their own product data, where can they find you, and what can you help them with?

Jay: Go to inriver.com — we’d be happy to help you understand your product data, where it’s serving you commercially, what the benefits are of having proper information management and information orchestration, and help you figure out how that’s impacting you in the real world. This isn’t an abstract problem — it’s about driving faster time to market and more revenue.

Floyd: Right on. And Willem, same to you — when someone realizes the data foundation is the problem, what’s the first move, and where do they find you?

Willem: Yeah, pivotree.com, or willem@pivotree.com by email — happy to talk to folks. We’re very low-pressure. We often do investigations, a lot of remediations for companies that have tried this but aren’t happy with the results, as well as net-new implementation support for solutions. So we do a lot of different things, all related to data services, as well as data platforms.

Floyd: And I’ll let our listeners know that if you’re looking for Klingon battle hymns, or you want to talk about Clippy, you can reach out to me — that’s probably less useful for you right now, so I’ll let these two shine and just give you their contact information. But that’s the data layer underneath the PIM question. Next time, we’re going to talk about the move itself — how do you actually get from where you are today to there, without stopping your whole business to do it? Matt will be back from a well-deserved vacation, and we’ll see you next time. Thanks.

Matt: Thanks for tuning in to this episode of Data Versus Commerce. New episodes drop weekly.

Floyd: So if you’re responsible for any part of how products get from a database to a doorstep, subscribe now on Apple, Spotify, or wherever you listen.