The robots.txt setting that’s hiding your catalog from ChatGPT
Matt asked ChatGPT to fix his look. He uploaded a photo, got some suggestions (glasses, which he rejected on the grounds that he’d end up looking like me, which I felt was pretty rude), then asked where to actually buy the stuff. It handed him a list of stores. Which raises the question now haunting retail teams: how did it pick those stores, and why those and not yours?
I’ve now spent months down this rabbit hole, and the first answer is almost embarrassing in how mundane it is. A lot of retailers are invisible to AI shopping tools because their own security settings are slamming the door on the crawlers. The robots.txt file, plus the anti-fraud and anti-DDoS rules that keep bad traffic out, are also turning away ChatGPT, Perplexity, and Gemini. Never get let in, and you’re not ranked low. You’re not in the room.
Once you’re inside, the model shops a lot like a person on a serious mission. What am I after, how fast can I get it, what’s a fair price. It checks the core product data first (whether you have this item in this color and this size), then price and availability, then trust signals like return policies and where the thing was made.
The newest layer, and the one growing fastest, is outside validation. The model doesn’t care about your brand story because it’s a machine. It’ll go looking for a GQ, a Runner’s World, a Reddit thread that vouches for you, the way you’d ask around before trusting a shop you’d never heard of.
This is where the marketers start to perk up, because brand isn’t dead and content still counts. It just gets consumed differently. The lush copy that makes a jacket feel like luxury means nothing to the AI model directly. The attributes underneath that copy mean everything.
“Soft,” “warm,” “puffy” have to become a fill rating and a temperature rating, so that when someone types “warm winter jacket that’s actually puffy,” your product reads as an answer. You can also train the model to look for you, through content and, increasingly, through ads, which are already creeping into these tools.
The traffic is growing fast, but it’s still small. Nobody’s moving millions through ChatGPT yet. What you’ve got is one more channel to feed—with its own quirks—stacked on top of the store, the site, mobile, and social. The consolation is that the on-site work, clean data and accurate availability and clear pricing, feeds your regular search rankings too. You’re paving a road everything drives on.
Measurement is still the ragged edge. Some emerging tools track how often you get recommended and which prompts surface you, and a few platforms now report AI referral traffic directly. There’s a hole in the map, though. When a shopper gets recommended, closes the app, and walks straight to your site, you often can’t see where they came from. What you can see is that these buyers convert far better than organic or paid search. They show up already sold.
So what do you check tomorrow? Start embarrassingly basic. Confirm you’re not blocking the crawlers. Make sure your product data is complete enough for a model to confirm you’ve got exactly what was asked for, down to the attributes and the UPCs and the GTINs. Round it out with the stuff that closes the sale, like price and availability and pickup and shipping policy. Then go find the third-party sources that should be name-checking you, and make sure they are. It’s early, and the race is barely off the line. The retailers who win it are the ones handling the boring basics before they chase anything clever.
Listen to Episode 4 of Data vs. Commerce wherever you get your podcasts.
