← Selected work PulseQueue / Case study Source ↗

PulseQueue

Retries are easy. Correct retries aren't.

Live · lease-safe retries · Updated 2026

TypeScript · React · SSE · Node.js

Hosted on Render. The first open may take up to about 60 seconds; the proof below remains available immediately.

Context
Independent queue-systems project · 2026
My role
Designer and sole engineer
Team
Solo build
Evidence
Lease-safe retries · visible dead letters

01 Problem

A retry that two workers both own is a double send. A retry with no backoff is a thundering herd. A retry that never dies is a poison pill.

02 Constraints

  1. 01A job is leased by at most one worker.
  2. 02Backoff is base * 2^(attempts-1).
  3. 03After maxAttempts the job is dead, not queued.
  4. 04The engine has no HTTP and no wall clock — time is injected.
  5. 05The dashboard is a subscriber, not the source of truth.

03 Architecture

Control path

04 Decisions

The queue engine is pure TypeScript. Clock injected.

Why Lease exclusivity, retry delay, and dead-lettering are unit tests, not a running server.

Tradeoff In-memory. The semantics are the demo; persistence is a store swap.

SSE, not polling.

Why Every state change is pushed. The UI never invents a tick.

Tradeoff Free hosts sleep. First event after wake can lag.

Dead-letter instead of infinite retry.

Why Poison messages must stop cycling so operators can see them.

Tradeoff Someone has to drain the DLQ. The dashboard makes that visible.

05 Failure

  • Two workers lease the same job.

    Lease is exclusive. Tests lock single-owner complete/fail.

  • A flaky job fails twice.

    Exponential backoff. It becomes eligible later, not immediately.

  • A job fails past maxAttempts.

    Marked dead. It leaves the active cycle.

06 Object

job_014 after maxAttempts

{ "id": "job_014", "type": "flaky", "payload": "partner-call-14", "status": "dead", "attempts": 3, "maxAttempts": 3, "workerId": null, "error": "partner timeout" }

07 Inspect

Live Open product ↗

Source GitHub ↗

Wake time Allow up to about 60 seconds on the first request.

08 What I would improve next

  1. Persist jobs and leases so ownership survives process restarts and multiple queue nodes.
  2. Add worker heartbeats and fencing tokens to prevent a paused worker from completing an expired lease.
  3. Publish saturation, retry-age, and dead-letter metrics with an audited redrive workflow.