Solution blueprint 01 · Cloud engineering

A store that stays open on sale day.

A growing brand's store slows and then fails in the first minutes of every sale, sells the last unit twice, and pays all year for the capacity it needs for one weekend. This is how we would rebuild it on AWS so it scales with the crowd and shrinks back after.

A blueprint, not a client story: a problem we see often in D2C retail, and how we would solve it.

The situation · what hurts

Where itbreaks today

The symptoms that bring a team to us, in their own words.

  • Down at the worst momentThe site slows, then fails, in the first minutes of a sale.
  • The last unit sells twiceTwo shoppers pay for one item; one order is cancelled after payment.
  • One server holds everythingThe app, the database and the images share one machine.
  • A bill sized for the peakCapacity for the busiest hour, paid for every hour of the year.
The solution · 10 parts, 9 flows

The blueprint,drawn

How the pieces fit together. Follow the numbers.

ShoppersWeb and appCFCloudFront + WAFCache and bot rulesECSStore serviceECS Fargate, auto-scalingPGAurora PostgreSQLWriter and read replicasS3S3Images and static pagesRDElastiCacheCatalogue and cartsSQSSQS order queueCheckout accepted at onceλOrder workersPaced to the databaseEBEventBridgeOrder eventsThe rest of the businessEmail, WhatsApp, warehouse123456789
  1. 1Shoppers → CloudFront + WAFEvery request meets the edge first
  2. 2CloudFront + WAF → Store serviceOnly what the cache cannot answer
  3. 3CloudFront + WAF → S3Images and pages from the edge
  4. 4Store service → ElastiCacheCatalogue and carts from memory
  5. 5Store service → Aurora PostgreSQLStock reserved inside the order
  6. 6Store service → SQS order queueCheckouts queued, never dropped
  7. 7SQS order queue → Order workersTaken at the pace the database can bear
  8. 8Order workers → EventBridgeOrder placed
  9. 9EventBridge → The rest of the businessEveryone downstream hears once
Blueprint 01 · D2C retail
The decisions · 4 that matter

What we woulddecide, and why

Reliable, scalable, secure and sensible on cost: the choices that get there, with the settings beside them.

01Scale

Scales with the crowd, not with hope.

The store runs as containers that add themselves as requests climb and step back as they fall, with an extra block switched on an hour before an announced sale. Catalogue pages and images are served from the edge, and browsing reads from replicas, so the primary database only takes orders.

The settings · scalespec
Compute
ECS Fargate on Graviton, 2 to 40 tasks
Scale on
Requests per target and CPU at 60%
Before a sale
A scheduled step up, an hour ahead
Cache
CloudFront for pages, Redis for carts
Database
Aurora: one writer, up to three readers
02Integrity

The last unit sells once.

Stock is reserved with a conditional update inside the same transaction as the order, so two shoppers cannot both win the last unit. Every checkout carries an idempotency key, payment confirmations are verified and claimed once, and an unpaid reservation is released by a delayed message.

The settings · integrityspec
Stock
Conditional decrement, never read-then-write
Checkout
An idempotency key per attempt
Payments
Signed webhooks, each event claimed once
Unpaid holds
Released by an SQS message after 15 min
Peaks
Orders queued, workers paced
03Security

Bots stopped at the edge.

Managed rules, bot control and rate limits sit in front of everything, so scalpers and scrapers never reach the store. The app runs in private subnets, secrets live in a vault rather than in code, and card data never touches the store at all because payment happens on the provider's page.

The settings · securityspec
Edge
AWS WAF managed rules and Bot Control
Limits
Per IP and per checkout
Network
Private subnets, no public database
Secrets
Secrets Manager, rotated
Cards
Hosted payment page, out of scope
04Cost

Pay for the peak only at the peak.

Everything above the everyday baseline exists only during the rush. The baseline runs on Graviton under a savings plan, the edge cache absorbs most reads, old images move to cheaper storage on their own, and a budget alarm speaks up the day spending drifts.

The settings · costspec
Baseline
Graviton, under a Compute Savings Plan
Peaks
Scale-in within minutes after the rush
Reads
Mostly answered by CloudFront
Storage
S3 Intelligent-Tiering
Guard
AWS Budgets and Cost Anomaly Detection
Design targets · agreed up front

What goodlooks like

Targets we would set together before the first line of work, and measure against after.

99.95%availability we design for
< 300 msp95 for catalogue pages
0units oversold, by construction
15 minbefore an unpaid hold lets go

The plan

  1. 01Baseline and load test today's store1 wk
  2. 02Edge, caching and WAF2 wks
  3. 03Containers and auto-scaling2 wks
  4. 04Order queue and stock integrity2 wks
  5. 05Load test at ten times normal traffic1 wk
  6. 06Cutover, runbook and handover1 wk
Next · blueprint 02 · Cloud engineeringLendingReady for the auditor, and for the worst day.

Sound likeyour problem?

Tell us what is breaking. We reply within a working day with how we would approach your version of it.