One in five retailers posted a product data role | EP. 12
Watch the podcast now
Episode 12 – One in five retailers posted a product data role- Transcript
Matt: Welcome to Data Versus Commerce, where we explore the messy middle between database and doorstep. I’m Matt Johnson.
Floyd: And I’m Floyd Blaikie. Let’s dig in.
Matt: Hey guys, welcome back to Data Versus Commerce. Today we’re talking about something that could potentially sound like just more AI noise, but I want you to hang with us, because what we’re going to talk about is the current reality in retail commerce — and it might actually be the difference between thriving and falling behind fast. The context for today’s discussion is the report we recently published, called When the Shopper Is a Machine. It tells the story of what we’re seeing in retail today — not what they’re saying they’re doing, but what they’re actually doing, who they’re hiring, what they’re building, and what’s keeping them up at night. To break this down, who better than our head of retail, Dan Ornstein? Dan, welcome back to Data Versus Commerce. For those who might have missed your first episode with us — if you did, go back and listen, it’s great, we talked about AI search. But Dan, give the listeners a quick intro — who you are and what you do.
Dan: Sure, thanks Matt, a pleasure to be here. The first time was so much fun, I had to come back — always a good discussion with yourself and Floyd, who unfortunately can’t make it today. My role at Pivotree is as head of our retail industry practice, which is our biggest segment — the vast majority of our revenue comes from work we do with retail customers across the globe. We serve major retailers in Europe, North America, and Asia. On one hand, it’s about excellent customer service for the work we do and making sure our clients are happy with what they’re getting. On the other, they really look to us to help figure out their strategy going forward around commerce — what are the major trends, how do they drive revenue through those channels and expand them, what do they need to watch out for in terms of customer expectations as shoppers figure out what to buy from whom, or technology changes we see that are inevitably disruptive in the commerce space. Even going back to where this whole space came from in the beginning — the internet showing up and creating a whole new channel no one would’ve thought of before. Obviously the topic of the day, as you mentioned, is AI as a disruptive technology in commerce, and the research we did in Q2 — we’ve started this process of producing a quarterly report on the state of retail — was a bit of a surprise, I think, when you look at the stories and hype that’s been out in the market around retail brands and what they’re doing on their commerce channels, versus the brass tacks of actually figuring out how to make those lofty objectives come to pass. I’m sure as we get into it, what I’m alluding to will become pretty clear pretty fast.
Matt: Yep, it will, and there’s so much to cover — we typically try to keep our episodes around 20 to 30 minutes at the most, and that’s way too short a time to really dig into all we put together in this report. So we’ll post a link in the comment section — if you’re seeing this on a social network, or if you’re on our website, dataversuscommerce.com, you’ll be able to access that resource there. But Dan, let me start with my first question, which is really what surprised me most in this report. We tracked over 1,600 different retailers in North America from April to June, and one in five of those retailers posted a product data role — far and above any other category of employee growth, more than AI roles, more than commerce platform roles, developers. This feels like a story that tells us something about where commerce operations are heading. What’s your take on the growth in product data roles?
Dan: I think what it tells us is that the smart retailers know where they have a bottleneck. While everybody’s kind of assuming the next big roles will be AI hires — creating agents of some sort — that work is kind of handled by engineering teams learning what’s new and taking care of it. So I don’t think we quite see that level of hiring except at senior levels. But what it really tends to shine a light on is the biggest constraint in making any of that work, which is all about data. AI engines, LLMs, all these tools out there, are massive consumers of information — the data underlying that. If your data isn’t up to date, real-time, accurate, and accessible through the integrations it needs to find your information, you can achieve nothing with AI. So the agents we see retailers bringing to market — whether it’s Kroger’s AI shopping assistant that helps you put together meal plans and fills your shopping cart — just think about what it’s looking for to do that for you. If you say, “I want a primarily vegetarian weekly budget and 10 recipes,” that’s all data it’s going to go through. It needs to know what products you’re selling, the inventory availability in my local store, pricing, if you give it a budget constraint, delivery information, and so on — all of that comes back to data. So it’s not that surprising that retailers are finding they have to add to the teams that manage that core information — all the appropriate attributes, and expanded new attributes that the LLMs are looking for, well beyond the typical ones that used to be there for SEO in e-commerce.
Matt: I’d be really interested to learn more about what — because obviously good enough is no longer good enough, right? That’s my takeaway. This isn’t like B2B commerce, the world I live in, where you’re just worried about the basics, Maslow’s hierarchy of needs, just trying to survive — you need food and shelter. In retail, we’re looking at how to go from good to great. Can you talk a bit about what great data looks like? What are these roles going to focus on — what’s the work they need to do today that maybe they haven’t done historically?
Dan: It’s not necessarily that they haven’t done it — I think the demands on it are higher in terms of accuracy and recency, meaning up to date to the minute or less. These are people and teams focused on owning the specification — the basic block that says, “This is a blue T-shirt of this size,” the core taxonomy. But then it’s everything else that helps LLMs and human shoppers figure out if this is the product they want — images, copy, descriptions, pricing, and the attributes that separate it from other products in your catalog or the competition. Things like the fabric it’s made of, its recyclability, sustainability-related information. Sourcing information can become important, not just for tariffs, but because people want to make sure their product is coming from somewhere for any number of reasons. It even goes beyond the core product information into what other teams manage, like reviews and recommendations — the agents are looking for all of that. What’s interesting, at the end of the day, is that all those visuals and information that we can consume as humans, and forgive if there’s an error — say we’re looking at a picture, looking for the gray version of something, and it’s mislabeled — it’s not really a problem, humans oversee that, we forgive the error if it’s a brand we want to buy from, we’ll figure out that blue also exists and buy blue. An LLM doesn’t care. If it’s looking for gray and your attribute list has it tagged wrong as blue, you don’t have a gray shirt as far as it’s concerned. It’s not going to recommend you, it’s not going to buy from you. Once we get there, you’re toast — and you may never know, because all of that’s happening on a surface that isn’t your store surface. It’s happening in ChatGPT, Perplexity, Gemini, Copilot, whatever I’m using. I may never know, as the human giving that instruction to the agent, that it even looked at this store to begin with, let alone skipped over it because it couldn’t find the right information to confirm this is the product I’m looking for.
Matt: That’s really fascinating, because a lot of people think AI in commerce and immediately think about the customer experience — putting together that shopping list, a tailored fit for you, whatever. But what we’re starting to see is that AI is actually being implemented more in the background of commerce operations than in the front-end customer experience. That makes a lot of sense to me, because that’s where all your friction is internally, that’s where your labor is. Besides Wish, not many brands are willing to put AI in front of their customers yet — the ones that are market-leading, we talked about some — Target, Walmart, Albertsons recently did this, Kohl’s just today, I saw an article about Kohl’s doing it. But on the back end, what’s happening there? Where are companies investing in AI in their commerce or data operations?
Dan: They’re investing in a few places. One is trying to gain an efficiency advantage, in terms of AI tools taking care of some of the work at the scale and pace it needs to be done, without having to create a gigantic team of people to do it — humans in the loop, of course, to make sure it works. But the ones who get it right are going to spend their time auditing their catalogs, fixing their taxonomy, reconciling pricing across channels — none of which is necessarily visible to anyone in the front end. But if you get it wrong, the front-end AI experience is terrible. If you’re making a promise to the market, to the customer, about how an AI tool is going to help them do something, and the data feeding it is incorrect, inaccurate, or out of date, you’re degrading the experience. Like anything digital, trust is the key thing — whether it’s online banking, the standard e-commerce experience we’ve all come to know, or now the evolving AI experience. If I’m working through your journey in an AI-enabled digital channel and something’s not right, or you can’t answer the prompt I’m giving you and I end up in a loop, I’m gone. I’m going to bail on it — and the danger is I bail on you entirely, not just default to some other way of doing it. I might say, “To hell with this, I’m going with your competitor,” because I’ll just go to their site and do it — I was experimenting with you. So it’s great if you build a nice storefront, but if it’s on top of the same cracked foundation you have today — those integrations behind the scenes that bring up price, availability, specific SKU availability, whether it’s online or in-store, how soon can I get it to you — any of that information that informs whether an LLM would recommend your product and your store to the consumer — if you don’t have that, you’re not going to be able to transact at scale, or at all, through your AI layer.
Matt: It’s crazy, the stakes are really high. We’re in this transition mode where we understand the potential, but smart retailers and distributors are understanding we need to ensure the foundation is solid, and invest there first. Now I’m going to change gears a bit, because one of the interesting sections in this report talked about the concept of the “consolidation tax.” That’s how you worded it — just as an example, we had massive mergers and acquisitions, DICK’S and Foot Locker, Ferrero and WK Kellogg, there are many more we could name. You called it the consolidation tax — can you walk us through what that means at the data layer? What’s actually happening behind the scenes that most people don’t realize — all the messy work that goes on to make these acquisitions really work?
Dan: Sure. Based on a number of years of experience around post-merger integration and putting two companies together, the headline is always, “This is going to be great, because customers from both banners or brands can now work together, we can get synergies from a supplier perspective for various categories and inputs, depending what kind of company I am.” But behind the scenes, that’s all about data. Who are the suppliers from both companies? Where are the opportunities to gain purchasing scale across certain ones of those? Who is the customer — are Matt and Dan both customers of both brands, or is it only Dan, and is it actually the same Dan that’s a customer across both companies? We certainly don’t want to confuse the customer with offerings from both brands they can’t execute in-store or online. Loyalty programs come into this in a big way — how do you merge your loyalty program, make sure that golden record of the customer is in fact the customer? All of this comes back to data. Am I selling the same product at a different price under my two different banners — do I want to do that, do I not want to do that? All of this represents the data behind the scenes that needs to be normalized, figured out for similarities, joined together, and potentially integrated in the flow of data across what we call the “every product moves twice” idea — the physical product might move from whatever DC to the customer’s house, but the data moves with it too. It takes a long time to actually achieve the synergies, or the benefits, to use the English word instead of the M&A word, from a large-scale post-merger integration. A lot of it is really hard work in process, technology, and functions to make sure there’s a consistent message to the customer. Brand identities may be maintained, but if there’s any intention of creating a single customer experience across banners, how do you bring those systems, information, and data together so it’s a unified, frictionless experience for the customer? Because as soon as the customer hits a problem — like, your Foot Locker loyalty card’s going to work at DICK’S Sporting Goods, except it doesn’t the first time you go to the point of sale, because it hasn’t actually been figured out yet — that becomes a real issue, even if you’ve managed to do it for 75% of the loyalty base but not the other 25%, and they’re still having a problem. It has to work for everyone out of the gate, otherwise it can’t be something you announce. So there’s a big tax that comes with mergers, which you typically see on the quarterly reports for public companies doing it, in terms of integration costs. It’s a big capital investment that has to go in to really get the benefit from that. Along the way, strategically, separate from what we’re talking about, you’ve got to make sure not to destroy the brand value you just tried to acquire. No easy task for the teams involved — usually a lot of time pressure to get it done as fast as possible. We’re seeing ways that AI-enabled technical delivery can help accelerate that, but it can only work at the speed of the people doing the work too. So there’s definitely a tax that hangs over an organization for a couple of years, especially when they’re really large.
Matt: I see this all the time in distribution — there’s a massive wave of consolidation happening across different industrial segments, and it’s difficult, to say the least, because as important as the brand experience is for retailers, it’s the pain that happens on the supply chain buy side of the business. The whole promise was that you’d leverage your ability to move product faster and consolidate warehousing, and all of that gets lost in what that actually represents in a database. So now you’re building that from scratch almost — companies going through these mergers and acquisitions, it’s not just, “Okay, now I’ve got two different operating systems and it’s a brand thing,” it’s way beyond that — it’s the integration of those into a master system, to have a record, both a supply chain record and a product record. So you’re basically starting over, aren’t you?
Dan: In some ways you are, and you have to. One of the things we’ve found, working with customers going through these kinds of consolidations, especially on the distribution side, is that unless you’ve got a real history of acquisitions and figured out what the industry calls a “playbook” — what you’re going to merge, what you’re going to leave separate — especially if you’re smaller distributors coming together, you don’t have the appetite for that kind of back-office, back-end integration across process, teams, and systems to really get the advantage. The required investment becomes too high, so they just leave the companies as they are, roll them up financially, and operate as they always did — becoming a financial rollup instead of a true integration where they could potentially get the benefits from sourcing or operational consolidation. Or it just takes them much, much longer than they originally intended, because they don’t have the stomach for it, or the investment required upfront — it’s just generally really expensive to try to do it early, so you end up doing it over time instead, which can work. There’s no right or wrong way to do this, but there’s a bit of expectations management involved — deciding, “This is the way we’re going to do it, now let’s go do it, we’re going to do it over three years, or two, or five” — and then managing toward those timelines appropriately, so whoever the stakeholders are, public or private, they’re along for the same ride, and you don’t get a misalignment, because then you run into a lot of pressure and short-bench challenges.
Matt: I love that, and I think you’re answering at least part of my last question here, as we turn the corner — because for people hearing themselves or their businesses in this report as we talk about it right now, we’ve talked about the integration layer, the importance of managing that consolidation tax with good integrations and foundational data work. They know data is more critical than ever, with all the AI tooling and AI search happening. My question is always this: you cannot boil the ocean, you cannot do everything, nobody does it perfectly — I don’t care if you’re Amazon or Walmart, nobody’s doing everything the right way. What’s the next right thing a retailer can do if they know they’re behind the ball on AI-ready data, and behind the ball on managing disparate and legacy systems? What are your tips for those listeners?
Dan: The tip I’d give our listeners is: run that assessment or audit before you do anything else. Pick your top 100 products, top 10 categories, however you define the core of the business you want to evaluate, and check whether the data that shows up in the different channels for those products is accurate, up to date, and consistent, wherever a machine’s going to read it. If it’s going to read it online in your store, that’s one thing — is it going to show up through the LLM models out there, the ChatGPTs and Perplexitys of the world? Is it going to show up in a social channel? Make sure it’s up to date, accurate, and consistent across those channels too, so you get an honest picture of where you actually stand — then you know the size of the gap, where it is, and you can prioritize fixing it so you’re actually ready for this agentic future. We’re already seeing it — we work with a number of retailers, and we’ve had this conversation with distributors too, who really want to understand whether they’re showing up in these models. They’re seeing traffic rise, visits rise, referred traffic rise from AI platforms — conversions are coming, but they’re definitely seeing traffic rise, and in many ways they’re trying to understand how and why they’re showing up, why they’re being recommended, and what prompts lead somebody to recommend their store over someone else’s, so they can influence the models. In order to influence the models, they’re quickly finding they have information gaps to fill — whether pure product gaps, third-party referrals and references that create the trust factor LLMs are evaluating, inventory information — not just yes/no availability, but very specific SKU-level, up-to-date inventory information — as well as return policies and other elements that create differentiation between brands. Everything may be completely right, but if my return policy is either absent or worse than my competitor’s, the LLM is going to drop me and go to the competitor, all else being equal. So that audit tells you where you stand in terms of what we call “agentic readiness” — we have a tool that can help companies figure this out relatively quickly, or at least get a snapshot of where the issue may lie. Is it in your product data source? Your pricing source? Your inventory or OMS? Is it just missing website content? That starts to pinpoint where to prioritize your efforts to fill those gaps.
Matt: I love that — you’re absolutely right, it’s, we’re data people, so use the data to determine the next right thing to fix the data. That seems logical to me. So where can folks get ahold of that assessment, Dan — is that something they email for, or can we drop a link?
Dan: It’s something they should probably go to the Pivotree website for and fill in the contact request, and we’ll run it for them and follow up with a report. Doesn’t take long.
Matt: Awesome. So guys, you’ve got the report — the Q2 retail report is live on the website, we’ll link to that for sure. And if you’re interested in Dan’s agentic commerce assessment tool, that’s going to be a lot of insight — let’s see, how much does that cost? Zero dollars, to get a good snapshot of where you’re at.
Dan: And then, depending on whether the company wants some support and help addressing those gaps, that’s a different aspect to it. But it definitely provides a first pass across about a dozen dimensions we’ve been talking about — a bit of scoring as to whether you’re poor, moderate, or advanced, which then highlights what areas you should tackle first.
Matt: Love that, love that — and then we’ll help you get into the deep dive. Guys, this has been awesome, Dan, always a pleasure. I’m going to wrap up here and say you’ll be back soon for sure — we’re going to do one of these every quarter, at the very least, so you’ll get to hear from Dan Ornstein once a quarter as we go through these reports. I’m really excited about it — we of course also did one for the State of Industrial Distribution. But just wrapping up, we talked about something that’s not the sexiest subject, but it’s the slow, foundational work that makes sure your product information is right, in the way your consumers are searching for your products today. I’d just say again, check out the link in the comment section, check out the link on the website, dataversuscommerce.com — you can go watch previous episodes there. One last thing I’ll say: I’ll be back for the next episode, and we’re starting to pick up some steam with this podcast, so if you’re a listener and you haven’t subscribed, and you haven’t dropped us a rating or review, please do that — it helps us tremendously as we reach out to more data commerce nerds just like you, me, and Dan. Guys, it’s been a pleasure, and we’ll catch you on the next episode.
Dan: Thanks, Matt. Take care.
Matt: Thanks for tuning in to this episode of Data Versus Commerce. New episodes drop weekly. So if you’re responsible for any part of how products get from a database to a doorstep, subscribe now on Apple, Spotify, or wherever you listen.