How to Design an eCommerce Platform for 10 Million+ SKUs With Arvyn

Ten million SKUs looks fine on a slide. It looks very different at 2 a.m., when a reindex job is still running and the merchandising team asks why half of yesterday’s price changes aren’t live yet.
Most catalogs don’t arrive at this size on purpose. A retailer acquires a competitor, or a manufacturer decides every bolt variant deserves its own listing, or a marketplace integration dumps three million SKUs into a system that was sized for 200,000.
Whatever the path, the architecture that worked fine a year ago starts failing in ways that are hard to explain to a board deck: reindexes that used to take twenty minutes now take six hours, inventory counts drift, category pages load like it’s 2009.
Below is what actually needs to change for faster ecommerce merchandising, not the feature checklist, once a catalog crosses into eight figures of SKUs.
The data model comes first
There’s a temptation to treat this as a search problem. It isn’t, really, not at the root. Search inherits whatever mess sits underneath it.
A flat product table with a couple of joined attribute tables works fine at small scale. At 10 million rows, especially once you add variants, kits, bundles, region-specific pricing, and category-specific attribute sets (a bolt and a dress do not share a schema, and pretending otherwise is how you end up with forty nullable columns), query plans start choking.
What most large catalogs actually look like under the hood is less “product table” and more “product graph” β parent items branching into variants, variants tied to multiple price books, availability spread across warehouses.
The fix is separating the system of record from the serving layer early. A PIM or ERP holds the raw data; a serving layer, shaped around how the storefront actually queries things rather than how the source system stores them, sits in front of it.
Serving storefront traffic straight off the system of record is probably the single most expensive architectural mistake at this scale, and one of the most common.
Search and facets need their own plan
Faceted navigation, letting someone filter by size, color, price, brand all at once, gets exponentially harder as SKU count climbs. The usual approach is an inverted index (Elasticsearch, Solr, whatever), and it works.
But it comes with a tax people underestimate: every catalog change requires reindexing, and at ten million SKUs that’s not a five-minute job. Hours, sometimes an overnight batch window.
That lag has a real, specific failure mode. A merchandiser marks something out of stock, and the storefront keeps showing it as available until the next index run finishes. Multiply by thousands of daily changes, and you get index drift: products that exist in the search layer but not really, in any sense that matters to a shopper who just tried to buy one.
Two responses to this, generally. Throw more infrastructure at the reindex: more shards, faster compute, which helps but never closes the gap fully. Or rethink whether an inverted index needs to be the source of truth for browse and facets at all.
Change data capture, streaming every catalog mutation into a flattened, query-ready read store, can answer facet counts directly, no rebuild required. This is one of the few spots where the architectural choice matters more than which vendor’s logo is on the box.
Worth a specific mention here. Arvyn, built by RBMSoft, was designed around exactly this problem β no reindex-and-wait cycle, catalog changes stream continuously into one read layer, and browse, facets, and merchant preview are all served from it without a separate index rebuild or nightly batch job sitting between a change and it going live.
For catalogs where “near real-time” is the actual requirement, not a nice phrase in a pitch deck, that removes a whole category of reconciliation work engineering teams otherwise carry forever. Worth a look at Arvyn if this is the exact wall you’re hitting.
Inventory and pricing as events, not syncs
A few minutes of inventory lag is tolerable at a small scale. At ten million SKUs spread across multiple fulfillment centers, marketplaces, and drop-ship vendors, that lag turns into overselling, and overselling at this volume isn’t a rare edge case. It’s a Tuesday.
Teams that handle this well stop thinking of inventory as something you sync periodically and start treating it as an event stream. Every stock movement, reservation, and cancellation gets published as an event, and the storefront’s availability view is a projection off that stream, current within seconds.
Pricing works the same way once you’ve got multiple price books, currencies, and negotiated B2B rates stacked on a base catalog β recalculate on read rather than baking a single price into the product record.
Caching has to assume it will fail
A catalog this size will never be fully cacheable, and pretending otherwise just produces caches that are stale half the time. What tends to work: category and listing pages cached with short TTLs and event-driven invalidation, product pages cached per-SKU with invalidation tied to the same catalog events, search results computed fresh off a fast denormalized layer rather than the raw source.
CDN edge caching helps a lot for static content. But the real bottleneck underneath the cache is usually the query pattern in the database, not the cache itself. Bolting on Redis without fixing that just gets you a faster path to the same slow answer.
Data quality is a full-time job, not a cleanup sprint
Missing images. Absent prices. Duplicate SKUs from a merged feed. Incomplete attribute sets. At this scale, these show up daily, not once a quarter, and treating them as occasional cleanup work rather than a structural problem is how catalogs end up with thousands of broken product pages nobody notices until a customer does.
Automated validation belongs in the ingestion pipeline itself, catching a missing price or category before it reaches a live page.
And merchandisers need to see what a change will actually look like on the storefront before it ships, not after β a live preview environment that mirrors production, rather than a staging cluster that drifts out of sync over time, saves an enormous amount of the back-and-forth that otherwise eats a team’s week.
Ranking gets harder, not easier, with more products
There’s a common assumption that more SKUs means better personalization, more signal to work with. In practice it’s closer to the opposite. Ranking a 200-item category page is a solved problem.
Ranking one with 40,000 variants across a dozen brands, each with different margins and different stock positions, is not something generic relevance handles well straight out of the box.
Retailers at this scale usually need ranking tunable toward a business outcome β margin, sell-through on aging inventory, pushing in-stock items ahead of ones sitting three states away. Whatever the goal, it has to run against live data, because a model trained on last week’s stock levels will happily promote something that sold out yesterday.
Independent scaling beats uniform scaling
A monolith that scales as one unit forces over-provisioning of the parts that don’t need it, just to keep pace with the parts that do. Search traffic, catalog ingestion, checkout, content delivery β these have completely different load curves and different failure tolerances.
Checkout needs to hold up during a spike. Ingestion can tolerate a short delay during a big supplier feed update. A slow ingestion job shouldn’t degrade page load for someone browsing right now, and a search outage definitely shouldn’t take checkout down with it.
This is the real argument for headless or composable architecture at this scale. Not because it’s the current trend, but because each piece can scale, and fail, on its own.
Governance, or who’s allowed to touch what
Nobody puts this on the architecture diagram, but at ten million SKUs it belongs there. Multiple teams, sometimes multiple companies if there’s a marketplace or supplier feed involved, are all writing to the same catalog.
Without clear ownership rules for who can override what, you get the same product with three conflicting descriptions depending on which feed last touched it, and a support team fielding complaints about a price that three different systems each think they own.
The catalogs that stay clean tend to have an explicit hierarchy of trust baked into the ingestion layer: which source wins on price, which wins on imagery, what happens when a manual merchandiser override collides with an automated feed update.
It’s not a glamorous problem to solve, and it rarely gets budget on its own, but skipping it is how a ten-million-SKU catalog slowly turns into ten million small arguments between systems.
Design for the catalog you’ll have, not the one you have
The most common mistake here is building for current volume instead of trajectory. A retailer at 3 million SKUs today, growing through supplier expansion, will hit 10 million faster than the original build accounted for.
Schema decisions and how tightly the storefront is coupled to the system of record are much harder to unwind later than to get right the first time.
The retailers who grow into this smoothly aren’t necessarily the ones with the biggest budgets. They’re the ones who treated the catalog as a real-time problem from day one, instead of retrofitting real-time features onto something batch-oriented after it started creaking. At ten million SKUs and climbing, that’s really the whole game.


