PORTFOLIO/02/Slipstream

3 screens · click to open live

Multi-cloud control planes metered down to the invoice

I helped grow a production control plane for deploying and securing application services from an established two-year-old platform into a 100-tenant system spanning six clouds. Over three and a half years, I built core capabilities across infrastructure, observability, security, and billing, but the work that demanded the most engineering discipline was metered billing: turning continuous consumption into an accurate, recoverable calculation that could withstand failure without ever producing the wrong invoice.

When I joined, the platform already had paying customers and the basic control plane was working, but much of the infrastructure still needed to be built. That meant the work had to happen around a live production system rather than behind one. I was one of three full-stack engineers, with a frontend developer joining in the final year, and ownership was deliberately end to end: each engineer was responsible for complete subsystems rather than working within a particular technical layer.

That model carried through the platform as it expanded. The time-series layer allowed it to report what infrastructure had actually done rather than what it had been configured to do, while provider integrations extended support across clouds and introduced node provisioning and live installation feedback so stalled deployments could be identified before they became failures. Alongside those systems, I built the WAF interface and migrated the existing frontend onto the new stack.

Billing became the most difficult of those subsystems because the requirement changed fundamentally. The existing implementation charged for discrete actions; the platform now needed to charge for continuous consumption, measured over time and enabled independently for each customer. What had previously been a lookup therefore became a calculation, and correctness had to hold even when the process doing that calculation did not.

A failed request, a partial calculation, a retry, missing telemetry, or a billing run that stopped halfway through could never result in a duplicate or incorrect charge. Failures had to be visible and recoverable, with the same inputs producing the same result when processing resumed. Six iterations appeared to work under normal conditions, but each was rejected because normal operation was not sufficient for a production billing system.

The seventh iteration finally satisfied the requirement and shipped. Since then, it has completed every billing cycle across every tenant without a single billing failure.

ROLE

Full-stack Engineer · Three-person team · Three and a half years

Metered billing and invoicing, time-series integration, and cloud-provider integrations. The interfaces over them, including provisioning, live deployment progress, and security policy. Migration of the existing frontend onto the new stack.

  • A change going out in waves rather than all at once, with each region's progress and health visible while it moves. The interesting state is not the one where everything succeeded. It is the region that stopped, and what the system did about it without being asked.

  • The conditions that halt a rollout, held as configuration rather than buried in the code that enforces them. A rule you can read is a rule you can argue with before it fires, which is most of the difference between a policy and a behaviour.

  • Rules in front of the same estate, some enforcing and some only watching. Turning one on is a decision with a blast radius, so monitor mode exists to say what enforcement would have done before anyone commits to it.

Reference implementation. The original is under NDA.