Deep Dives

How AI Catalog Scoring Works at ReFiBuy

By Scot Wingo
Back to Catalog Insights
How AI Catalog Scoring Works at ReFiBuy

When we talk to catalog managers about AI agent visibility, the conversation usually hits the same wall: "How do you actually measure this?" It is a fair question. Organic search visibility has years of tooling behind it. AI agent visibility does not. This post explains the 12-point scoring system we built at ReFiBuy, what each dimension measures, and how we use the scores to decide which listings to fix first.

Why a Score at All

Before explaining what we measure, it is worth explaining why a numerical score makes sense here. An AI shopping agent does not make binary decisions. It does not decide "this product is qualified" or "it is not." It constructs a response that incorporates the products it can most completely and accurately characterize. A product with a score of 85 does not get recommended at ten times the rate of a product scoring 8.5; the relationship is not linear. But a listing that scores below a rough threshold on multiple dimensions tends to be absent from agent responses entirely, while listings above threshold appear with usable specificity.

The 12-point system was designed to be independent dimensions, not a waterfall. Failing one dimension does not cascade into failure on others. This matters for remediation: when you know a listing scores well on attribute completeness but poorly on description agent-readability, you know exactly what to fix without disturbing what is working.

The 12 Scoring Dimensions

1. Attribute Completeness

Does the listing include the key structured attributes for its product category? Attribute requirements vary by category. A laptop needs processor, RAM, storage, screen size, and operating system. A sleeping bag needs temperature rating, fill type, fill weight, and packed dimensions. We maintain category-level attribute requirement maps and score against those. A listing with three of six required attributes for its category scores proportionally.

2. Attribute Format Normalization

This is distinct from completeness. A listing might have weight listed in three different places with three different unit formats: "2.3lbs", "1.04 kg", "approximately one kilogram." An AI agent trying to compare weights across products in a category has to resolve that ambiguity. Normalized attributes in a consistent format score higher.

3. Description Agent-Readability

We score descriptions based on whether they answer evaluative questions rather than just list features. An evaluative description says who the product is for and what problem it solves in a specific context. A feature list says it has a 400D nylon shell. Both can be true, but the evaluative framing is what agents use to match products to natural language queries. We use a heuristic model that looks for use-case language, comparative qualifiers, and scenario specificity.

4. Title Specificity

Product titles optimized for keyword search often front-load category keywords and leave the specific product identifier in the middle or end. AI agents want to understand from the title what the product is and how it differs from similar products. A title like "Outdoor Research Ferrosi Hooded Jacket Men's Softshell" is more agent-friendly than "Men's Softshell Jacket Outdoor Wind Resistant Hooded Trail Hiking." Both might perform similarly on keyword search. The first tells an agent significantly more about the product.

5. Price Clarity

Is the current price present in machine-parseable format with currency? This sounds basic but fails surprisingly often. Sale prices without regular price context, prices embedded in JavaScript rendering, and prices that require interaction to reveal all score lower. Schema markup with the correct offers.price and offers.priceCurrency fields passes this check automatically.

6. Availability Signal Clarity

Similar to price clarity. AI agents hedging on product recommendations often do so because availability is ambiguous. A listing that clearly indicates in-stock, out-of-stock, or pre-order in both the page content and schema markup scores higher than one where availability requires inference.

7. Schema Markup Coverage

We score schema in two sub-dimensions: presence (is there a Product schema object on the page) and field coverage (which fields are populated). Presence of a schema object without meaningful field coverage is worth less than it appears. The fields that matter most for AI agent use are name, description, offers, brand, image, and at minimum three to five category-relevant itemProperties.

8. Image Signal

AI agents increasingly use image data in product evaluation. Our scoring here is limited to whether structured image metadata is present and whether multiple images exist. We do not evaluate image quality directly, but the metadata signal contributes to how agents assess product legitimacy.

9. Review Signal

Aggregate review count and rating in schema markup. AI agents use this as a proxy for product credibility. A product with no review signal and a product with 200 reviews at 4.3 stars are not equivalent in agent reasoning, even if the catalog description is identical.

10. Brand Entity Clarity

Is the brand clearly identified in both the listing content and schema? Brand disambiguation matters when agents are comparing products from multiple sellers. A clear brand entity in the schema reduces ambiguity about provenance.

11. Cross-Platform Consistency

For retailers syndicating to Google Shopping, Facebook Catalog, or other feeds: does the content in those feeds match the on-page content? Inconsistency between the product page and its feed entry is a negative signal for agent trust. We check this where feed access is available.

12. Return and Shipping Clarity

Shipping and return terms in machine-readable format. This is one of the newer additions to our scoring model. As AI agents become more sophisticated in purchase flow, the ability to surface shipping cost and return policy directly in a recommendation becomes relevant.

How We Prioritize Which Listings to Fix

A score alone does not tell you where to invest effort. We apply two additional filters to produce a prioritized fix list. First, estimated traffic weight: higher-traffic and higher-margin products get elevated priority. Second, remediation difficulty: a listing that fails only on attribute format normalization is much faster to fix than one that fails on description agent-readability, attribute completeness, and schema coverage simultaneously. We surface the high-impact, lower-effort fixes first, then the high-impact, higher-effort ones.

The goal is not to achieve a perfect 12/12 score on every listing. That is an unrealistic standard and not necessarily the right investment. It is to bring every listing above the practical threshold on the dimensions that AI agents weight most heavily for its category, and to fix the highest-traffic listings first so you see real improvement in agent recommendation rates before you have touched every SKU in the catalog.

What the Scores Do Not Measure

The scoring system does not evaluate whether your product is actually good, competitively priced, or relevant to user intent at the query level. It measures whether the listing gives AI agents enough to work with. A perfectly scored listing for an overpriced product will still lose recommendations to a better-priced competitor. We want to be clear about that scope boundary: ReFiBuy gets your catalog into the evaluation set. What happens inside that evaluation depends on factors beyond listing quality.

Is your catalog ready for AI shopping agents?

ReFiBuy audits your product feed and rewrites the listings AI agents cannot evaluate. See how it works.

See plans

More from Catalog Insights