Catalog Optimization

Structured Data vs. Natural Language: What AI Agents Need

By Scot Wingo
Back to Catalog Insights
Structured Data vs. Natural Language: What AI Agents Need

There is a persistent debate in catalog optimization circles about whether AI shopping agents prefer structured data (attribute tables, schema fields) or natural language (well-written product descriptions). The answer is not one or the other. It is a specific combination, and most catalog tools were never designed to produce it correctly.

What Structured Data Does Well

Structured data excels at answering quantitative and categorical queries. When a shopper asks "show me tents under 3 pounds with a capacity for two people," the agent needs machine-parseable values for weight and capacity. A prose description that says "our lightweight two-person tent is ideal for backpacking" contains the same information, but the agent has to parse "lightweight" as a weight class (how light exactly?), infer the capacity from "two-person," and verify that "backpacking" implies a weight consistent with the query's 3-pound threshold. That inference chain introduces error.

Structured data eliminates the inference step. A weight attribute with a value of 2.4 pounds and a capacity attribute with a value of 2 persons give the agent direct answers without parsing. For comparison queries, structured data is even more important: "compare these five tents by weight and packed dimensions" requires structured values for all products to construct a clean comparison table. Agents cannot reliably compare products on attributes that are only implied in prose.

The categories where structured data matters most:

  • Physical dimensions and weight (especially when units need to be consistent across products)
  • Technical specifications with defined categories (processor generation, material standard, compatibility standard)
  • Regulatory or certification attributes (temperature ratings, safety certifications, material standards)
  • Price, availability, and shipping terms
  • Variant-specific properties (color, size, configuration)

What Natural Language Does Well

Natural language description is the primary mechanism for communicating use-case suitability, context fit, and differentiation reasoning. These are things that structured data struggles to express.

A structured attribute for "use case" might be a multi-select field with options like "hiking," "backpacking," "mountaineering," "car camping." That tells an agent this tent is suitable for those activities. A natural language description that says "designed for alpine climbing teams who need fast pitching on exposed ridges with minimal guying points" tells an agent something different: this is a specialized product for technical terrain, not a general backpacking tent. The distinction matters for matching to intent-specific queries, and structured multi-select attributes cannot capture it with the same nuance.

Natural language also handles comparative differentiation better than structured data. "Unlike most tents in this weight class, the design prioritizes internal volume over packed size, making it better suited for trips where sleep quality matters more than pack compression" is a statement that structured data cannot encode in any standard schema field. But it is exactly the kind of reasoning an AI agent uses to recommend this product over a lighter alternative to a shopper who expressed comfort preferences in their query.

The Combination Requirement

AI shopping agents use structured data and natural language in sequence, not as alternatives. The typical evaluation pattern works roughly like this: structured data handles the initial filtering (is this product in the right category, price range, and specification envelope for the query?), and natural language handles the recommendation reasoning (why is this specific product the right choice within the eligible set?).

A product that passes the structured data filter but has a weak natural language description will appear in filtered result sets but receive low recommendation confidence. An agent comparing five products that all meet a weight and price threshold will favor the one whose description provides the clearest explanation of use-case fit.

A product with an excellent natural language description but missing structured data will fail the filtering step and never reach the recommendation phase for specification-driven queries. It might appear in broader queries where no filtering is applied, but those are typically lower-conversion queries.

The practical implication: both layers need to be working. Structured attribute coverage is the prerequisite for entering the evaluation pool. Natural language description quality is the determining factor within the pool. Investing in one without the other produces partial results.

Where Most Catalog Tools Fail This

Most catalog management and PIM tools were designed around the Google Shopping feed model. They are good at managing attribute fields like title, description, product type, GTIN, brand, and price. They typically have limited or no support for the additionalProperty fields in Product schema that carry category-specific specifications, and they rarely provide tooling for evaluating whether prose descriptions answer evaluative questions rather than just checking for minimum word count.

The result is catalogs that are well-maintained by legacy standards and missing what AI agents need. The feed is complete, the prices are right, the availability flags are accurate, and there is still no weight attribute for apparel products (just "lightweight" in the description), no compatibility specification for electronic accessories, and no use-case anchor in the description for any of the outdoor gear.

We built ReFiBuy specifically to address this gap: to audit what is missing on both the structured data and natural language sides, and to produce the combination that AI agents actually need. Existing catalog tools will improve over time as AI agent compatibility becomes a standard requirement. In the meantime, the gap exists, and it is large enough in most catalogs to significantly affect AI agent visibility.

A Note on Balance

We want to be direct about one thing: stuffing structured attributes without improving natural language description produces marginal gains. Adding 20 additionalProperty fields to a listing that still has a keyword-stacked, benefit-claim-heavy description will get you past the filtering stage more consistently, but the recommendation quality you receive from agents will still be limited by the description quality. The structured and natural language layers need to improve together for the full benefit.

Similarly, rewriting descriptions to be beautifully evaluative and specific while leaving structured attributes empty produces similarly incomplete results. The agents that rely primarily on prose parsing can benefit, but the agents that use schema-first indexing will not. The combination is the target, not an optimization of one axis at the expense of the other.

Is your catalog ready for AI shopping agents?

ReFiBuy audits your product feed and rewrites the listings AI agents cannot evaluate. See how it works.

See plans

More from Catalog Insights