Skip to content
Runner Blog
Esc
navigateopen⌘Jpreview
On this page
Ecommerce Operations13 min read

Site Search for Ecommerce: A 60-Minute Audit

A practical guide to auditing ecommerce search queries, product data, ranking, filters, zero-result recovery, and weekly improvements.

A cream-and-terracotta editorial illustration of a storefront search workbench with product cards, filter controls, diagnostic paths, and a scorecard

Effective site search for ecommerce helps shoppers reach a relevant, purchasable product without translating their intent into the store’s catalog language. Improve it by first running a controlled query audit that separates product-data failures, retrieval and ranking failures, and search UX failures. The diagnosis shows what to fix, who owns it, and whether technology is the constraint.

Key Takeaways

  • Test a representative set of real query types, not only popular keywords or exact product names.
  • Score whether suitable products were retrieved, ranked visibly, explained clearly, and made easy to refine.
  • Trace each failure to product data, retrieval, ranking, or UX before choosing a remedy.
  • Treat zero-result pages as recoverable states with specific alternatives, not dead ends.
  • Baseline search behavior by device and cohort, then improve one failure class at a time.

What should ecommerce site search do?

Ecommerce site search should interpret a shopper’s words, retrieve plausible matches, order the strongest options first, and provide enough context and controls to continue. A result count does not prove success: irrelevant items create noise, and the right product buried deep in results is functionally hidden.

Treat search as four connected systems:

  1. Product data supplies titles, identifiers, categories, attributes, compatibility, use cases, price, and availability.
  2. Retrieval decides which products qualify for the result set.
  3. Ranking decides which qualifying products appear first.
  4. Search UX explains the interpretation and lets shoppers refine, recover, or choose.

This model prevents changing ranking weights when a required attribute was never stored, or redesigning filters when retrieval failed. Strong product data foundations and disciplined catalog management are part of search quality.

How do you audit site search for ecommerce in 60 minutes?

Run a timed audit with a fixed query set, an incognito browser, a mobile viewport, and a simple scorecard. Find reproducible failure patterns and record what a shopper sees before anyone explains the system’s intended behavior.

Minutes 0-10: Build the test set

Choose 12 to 20 queries covering the catalog’s main demand patterns. Use actual internal search terms when available, then add edge cases. Include in-stock, out-of-stock, new, and variant products. Do not select only known successes.

Baymard Institute’s ecommerce search research identifies eight useful query types: exact, product type, feature, use case, abbreviation or symbol, compatibility, symptom, and non-product. Its query-type research provides definitions and examples for each type (retrieved August 29, 2026). Use the taxonomy as coverage guidance, then adapt the words to your own products and customer language.

Query type Example for a hypothetical catalog What it tests
Exact AeroPress Clear or a known SKU Titles, model numbers, identifiers, aliases
Product type coffee grinder Category language and broad retrieval
Feature stainless steel burr grinder Structured attributes and multi-term matching
Use case grinder for travel Use-case tags, synonyms, and interpretation
Abbreviation or symbol 12 oz mug and 12-ounce mug Normalization of units and equivalent forms
Compatibility filters for AeroPress Clear Product relationships and strict compatibility
Symptom or problem coffee tastes bitter Problem-to-solution mappings or guidance
Non-product return policy Help content and whole-site routing
Typo cofee grindr Typo tolerance without irrelevant expansion
Commercial constraint grinder under $100 Price interpretation and visible refinement

Add a singular, misspelled, alternate, or differently formatted variation for important queries to expose brittle matching rules.

Minutes 10-30: Run and score every query

Enter each query exactly as written. Do not use autocomplete for the first pass because suggestions can hide weaknesses in submitted-query handling. Capture the top results, result count, visible filters, applied interpretation, and recovery options. Then repeat a smaller high-priority subset on mobile.

Score five dimensions from 0 to 2:

Dimension 0 1 2
Retrieval No suitable product retrieved Some suitable products, major omissions Expected suitable products are present
Top-result relevance Top results are wrong Mixed relevance or right items buried Strong choices appear first
Result clarity Cards hide decision facts Some useful facts are visible Key matching facts are easy to compare
Refinement Filters are absent or misleading Generic filters provide limited help Relevant filters and sort options support the query
Recovery Dead end or unexplained result Generic alternatives Specific correction, category, content, or related-query path

The maximum is 10 per query, but it is not an industry benchmark. Compare only your own query classes, devices, releases, and repeated audits. Keep raw notes because equal totals can require different fixes.

Minutes 30-45: Diagnose the root cause

Assign one primary failure class to every dimension scored 0 or 1. If ownership is unclear, inspect a known product’s indexed record before changing settings.

Failure class Evidence to look for Likely owner First corrective action
Product data Expected attribute, synonym, identifier, relationship, or stock state is absent or wrong Catalog or merchandising Correct the source field and reindex
Retrieval The indexed product has the terms or mapped attributes but never enters the result set Search engineering or vendor Inspect tokenization, synonyms, filters, and query parsing
Ranking Relevant products are retrieved but buried below weaker matches Search and merchandising Review field weights, exact-match boosts, popularity signals, and demotions
UX Good results exist but cards, filters, labels, or mobile controls hide the path Product or design Expose matching facts and make refinement legible

Google Merchant Center’s product data specification is written for Google’s surfaces, but its data-quality principle is useful here: accurate titles, descriptions, identifiers, variants, price, and availability support matching and prevent eligibility or display problems (retrieved August 29, 2026). Your internal index may use different fields, yet it still cannot retrieve an attribute or relationship that the catalog does not represent.

Minutes 45-60: Prioritize the next fixes

Prioritize by query importance, severity, frequency, and fix scope. Start with failures affecting a meaningful query family with a clear cause. One missing field across a category may matter more than dozens of one-off synonyms.

Create a compact backlog with these fields:

  • Query or query family
  • Device and customer cohort observed
  • Expected behavior and actual behavior
  • Score by dimension
  • Root-cause class and supporting evidence
  • Owner and proposed change
  • Validation query set
  • Measurement window after release

Do not combine data cleanup, retrieval expansion, ranking changes, and interface redesign into one release unless they are inseparable. A narrow change makes the next audit interpretable.

How should search results help shoppers decide and recover?

Search results should make the engine’s interpretation visible, help shoppers narrow plausible matches, and provide an honest next step when nothing qualifies. Ranking, filters, and zero-result recovery solve different parts of that job, so test each one separately before changing them together.

How do filters and ranking work together?

Ranking should order plausible products, while filters express hard constraints and show how the query was interpreted. If blue linen shirt becomes product type, color, and material, showing those applied attributes makes the interpretation visible and reversible.

Useful filters depend on context: size and fit for apparel, dimensions for appliances, or compatibility for replacement parts. Generic filters repeated across categories create noise.

Audit filters with four checks:

  1. Availability: Can a shopper refine by the attributes expressed in the query?
  2. Accuracy: Do counts and values reflect the current result set and variant state?
  3. Persistence: Does changing sort, pagination, or device orientation preserve the query and selected filters?
  4. Visibility: On mobile, can shoppers see that filters are applied and remove them without reopening several panels?

Ranking inputs need a documented purpose. Exact identifiers, product names, text relevance, category fit, availability, merchandising rules, and behavioral signals may contribute. Popularity should not bury a clearly named feature or compatible model.

How should zero-result search recover?

A zero-result page should preserve the query, explain what happened in plain language, and offer the closest safe next actions. Recovery is not the same as filling the page with loosely related products. If no item satisfies a compatibility or safety constraint, an honest empty state is better than a misleading match.

Use the cause to choose the recovery path:

Zero-result cause Recovery action
Likely typo Show a specific corrected query and allow the original query to remain accessible
Over-constrained query Identify removable constraints or show which term caused the empty set
Known synonym gap Route to the equivalent category or rerun with the mapped term
Product unavailable Show the relevant category, compatible alternatives, or restock path when accurate
Non-product intent Link directly to the relevant policy, help page, or contact path
Unsupported problem query Offer related categories or a buying guide without pretending they are exact matches

Log the original query and recovery action. A zero-result rate can fall if the engine returns irrelevant products for everything, so sample recovered queries for credible paths rather than nonzero counts.

How should ecommerce search be measured?

Measure search as a sequence from query submission to useful engagement and purchase, using explicit definitions and your own baseline. Compare periods only when event collection, catalog availability, and query classification are stable. Segment by device, new versus returning visitors, market, and important query type where sample size permits.

Measure Definition Diagnostic use
Search usage rate sessions with a submitted search / eligible sessions Shows where shoppers choose search; not inherently good or bad
Zero-result rate searches with zero returned results / submitted searches Finds retrieval gaps, but requires relevance review
Search refinement rate searches followed by query edit or filter change / submitted searches Indicates exploration or friction depending on the query
Result click-through rate searches with a result click / submitted searches Shows whether results invite a plausible next step
Search exit rate search-result views ending the session / search-result views Flags dead ends while requiring page and intent context
Search-assisted conversion rate converting sessions with search / sessions with search Describes a cohort; it does not prove search caused conversion
Query success sample audited queries meeting the team's defined pass condition / audited queries Tracks controlled quality across releases

Google Analytics enhanced measurement can emit view_search_results when a results-page URL contains a recognized query parameter. By default, GA4 looks for q, s, search, query, or keyword, allows other parameters to be configured, and populates search_term. Google’s enhanced measurement documentation describes this behavior (retrieved August 29, 2026).

That event is only a starting point. It does not report relevance, returned products, zero results, filters, or asynchronous result changes without a qualifying URL. Validate it and add only the custom instrumentation needed for diagnosis. For broader cohort and funnel design, use the store analytics guide.

Do not adopt a universal target for zero results, click-through, or conversion. Catalog shape, customer intent, seasonality, and device mix change the expected values. Establish a baseline, inspect query samples, and compare equivalent cohorts and devices after a controlled change.

What should the weekly search iteration loop include?

The weekly loop should connect observed queries to one diagnosed failure class, one owned fix, and one validation step. The cadence matters more than a large quarterly redesign because catalog language, inventory, and customer demand keep changing.

Use this sequence:

  1. Review: Sample high-volume queries, zero-result queries, high-exit queries, and recently changed catalog areas.
  2. Classify: Assign query type and root-cause class. Keep an unknown state rather than guessing.
  3. Prioritize: Choose a small set based on importance, severity, recurrence, and fix reach.
  4. Change: Update the authoritative product field, retrieval rule, ranking rule, or UX surface.
  5. Validate: Rerun the affected queries plus neighboring queries that could regress.
  6. Measure: Compare the same device and cohort baseline over an appropriate window.
  7. Record: Keep the before-and-after evidence, owner, release date, and decision to retain or revert.

Maintain a regression set for exact products, important categories, attribute combinations, compatibility rules, and non-product intents. Rotate a smaller exploratory set from recent search logs.

If audits find missing product fields, repair the catalog contract before buying search software. If clean indexed records fail across query families, inspect the engine. If retrieval works but ordering or interaction fails, focus on ranking or UX.

Frequently asked questions

Good site search for ecommerce retrieves plausible products for the shopper’s language, ranks the strongest matches first, exposes the facts needed to compare them, and offers clear refinement or recovery. Judge it across representative query types rather than by visual polish or result count alone.

How many queries should a site search audit include?

Start with 12 to 20 carefully selected queries for a 60-minute audit, covering major query types, products, variants, stock states, and devices. Expand the permanent regression set as the team finds important failures; representativeness matters more than an arbitrary large count.

Should ecommerce search show out-of-stock products?

Show or hide out-of-stock products according to shopper intent, replenishment expectations, and available alternatives, but make the state explicit. An exact-model search may benefit from showing the product with a restock or alternative path, while broad category results may reasonably demote unavailable items.

When should a store replace its search technology?

Replace or add search technology when controlled tests show that clean, correctly indexed product data still cannot support required retrieval, ranking, scale, or operational controls. Do not use a vendor change to avoid fixing catalog gaps, ownership, or an unclear search interface.

How often should ecommerce search be audited?

Review search signals and a rotating query sample weekly, rerun the stable regression set after material catalog or search changes, and perform a broader audit when seasonality or product mix shifts. Choose a cadence the team can sustain and compare against its own baseline.

For more practical ecommerce operating guides, browse the Runner AI blog.

Sources

Last updated on August 29, 2026

Was this page helpful?