In October 2025 we ran the ReFiBuy rewriting engine on the live catalogs of three early-access retailers. We gave each of them 60 days of active rewriting and monitoring against a baseline period. Here is what we saw, what surprised us, and what it means for catalog teams thinking about AI agent visibility.
The Setup
All three pilot retailers were mid-size operations with established catalogs: between 800 and 4,200 SKUs each. Two used Shopify, one used WooCommerce. All three had meaningful Google Shopping traffic and well-maintained feeds by standard metrics. None had deliberately optimized for AI agent visibility before working with us.
We established a baseline by querying Perplexity Shopping and ChatGPT Shopping with product-category queries relevant to each retailer's catalog during the two weeks before rewriting began. We recorded how many of their products appeared in AI-generated recommendations and what form those recommendations took (cited with attributes, mentioned without specifics, or not present). We repeated the same query set at 30 days and 60 days post-rewrite.
We are not going to claim this is a controlled scientific study. We had three retailers, not three hundred. There are confounding variables we cannot fully isolate, including seasonal demand shifts and AI agent updates that may have affected baseline behavior. What we can report is directional signal from real catalog data, which is more useful at this stage than waiting for perfect methodology.
What Changed at 30 Days
The first thing we noticed at 30 days was not a dramatic jump in product recommendations. It was a change in recommendation quality for the products that did appear. Before rewriting, when AI agents cited these retailers' products, the citations were often vague: a product name and price, with no attribute-level specifics. After rewriting, agents were more likely to pull specific attribute data into their response, comparing our retailers' products on concrete dimensions like weight, material, or compatibility.
That matters because AI shopping agents have a qualitative filter as well as a recall filter. A product that appears in a response without supporting attributes is unlikely to be acted on. A product that appears with specific specifications that match the shopper's stated requirements is a fundamentally different kind of recommendation.
At 30 days, across the three retailers, the percentage of their rewritten listings that appeared in at least one AI agent response to a relevant query increased from a baseline average of roughly 12 percent to approximately 41 percent. This is not the same as revenue, but it is the prerequisite for revenue.
What Changed at 60 Days
By 60 days the pattern was more consistent. Retailers one and three showed similar trajectories. Retailer two, the WooCommerce merchant, showed a more modest improvement, and we think we understand why: their schema markup was inconsistent in ways that made it harder for agents to parse the rewritten content even when the descriptions themselves were good. The lesson there is that description quality and markup quality need to improve together.
Across the three retailers, AI agent recommendation rates for rewritten listings were approximately three times higher at 60 days than at the pre-pilot baseline. We define "recommendation rate" as the fraction of rewritten listings that appeared in agent-generated responses to category queries we tracked. The absolute numbers are small given the pilot scale, but the direction is clear.
One retailer (the outdoor gear merchant, Shopify, about 1,200 SKUs) shared that they had started tracking direct referral clicks from Perplexity Shopping in their analytics, a signal they had seen essentially zero of before the pilot. At 60 days they were recording a small but measurable stream of referral sessions from Perplexity. We cannot attribute all of that to the catalog work because other variables changed, but the correlation was strong enough to be interesting.
What We Got Wrong
We expected attribute completeness to be the dominant driver of improvement. It was important but not the only thing. Description agent-readability, meaning how well the product description answers evaluative questions a shopper would ask, turned out to matter as much for recommendation quality as attribute completeness did for recommendation frequency.
We also underestimated the schema markup dependency for the WooCommerce retailer. We are not saying that Shopify is better for AI visibility than WooCommerce; it is not a platform-level difference. It is a specific schema implementation difference that affected how agents indexed the rewritten content.
The third thing we got wrong was the expectation of a linear progression. Improvement was not smooth. At certain points during the 60 days, AI agent behavior seemed to update in ways that caused short dips in some metrics. That is not a failure of the catalog work; it reflects the underlying instability of agent behavior as these systems iterate. Part of what ReFiBuy has to do is track those shifts and re-score listings when agent criteria change.
What We Tell Retailers Going In
Based on the pilots, here is what we now tell retailers before they start. First, the first 30 days are about building the foundation: getting attribute coverage right, normalizing units, fixing schema. Expect modest gains. Second, the 30 to 60 day window is where you see the compound effect of having a larger share of your catalog properly formatted. Third, schema and description quality need to be addressed in parallel; doing one without the other limits your results.
We are not promising any retailer a specific outcome. Pilot results at three retailers are not a projection for your catalog. What we can say is that every retailer in the pilot started with essentially zero AI agent attribution and ended with a measurable and growing share. The baseline is low enough that the direction of improvement is not surprising. The question is always how big the gap is between your current catalog state and what AI agents need.
What Comes Next
We are opening the platform to additional retailers with Shopify and WooCommerce integrations. The core audit and rewriting engine that ran the pilots is the same system we are productizing. The main difference from pilot to product is the monitoring layer: rather than us manually tracking query results, the platform will surface listings that drop below threshold as agent standards evolve.
If your catalog currently gets little to no referral traffic from Perplexity or ChatGPT Shopping, the pilot results suggest there is probably something structural your listings are missing. The audit will tell you exactly what.
Is your catalog ready for AI shopping agents?
ReFiBuy audits your product feed and rewrites the listings AI agents cannot evaluate. See how it works.
See plans