Platform Architecture1 min read
Database sharding for authorization pipelines
Sharding a database is often the wrong first move. When it isn't, the choice of shard key is the decision you can't undo.
Sharding a database is often the wrong first move. Query optimization, better indexing, or read replicas usually give more runway. But past a certain scale — measured in transactions per second, not in database size — sharding becomes necessary. And the shard key choice is the decision you can't undo without a large migration project.
Written May 2025 from a scaling review.
The candidates
- Shard by transaction ID. Uniform distribution, but every read for a merchant's transactions hits every shard.
- Shard by merchant. Merchant queries are cheap; large merchants create hot shards.
- Shard by acquirer. Acquirer queries are cheap; acquirer skew creates hot shards.
- Shard by hash of merchant + time bucket. Balanced load; more complex query planning.
There's no universally correct answer. The right key depends on your query pattern.
What to measure before choosing
- The distribution of transactions per merchant. If a small number of merchants dominate volume, merchant sharding has hot-spot risk.
- The most common query shape. If most reads are "show me this merchant's last 30 days", merchant sharding wins. If most reads are "look up this transaction ID", hash sharding wins.
- The write pattern. Sequential ID inserts hit one shard at a time; hashed IDs distribute writes.
What to plan for
- Cross-shard queries will exist. Reconciliation, reporting, fraud analytics all need cross-shard patterns. Plan the query layer to handle it.
- Shard rebalancing will be needed. Traffic patterns shift; a shard that was fine on day one is hot a year later. Build the tooling to move data between shards without downtime.
- Backups per shard. Backup size stays manageable, but you now have N backups to test restore against, not one.
Sharding is a scaling tool, not an architecture. Reach for it when you need it, not before.