Series B SaaS DevOps company (anonymized)
Building a multi-tenant CI/CD platform from scratch
Designed and shipped the first version of their multi-tenant CI/CD platform in 14 weeks. Pipeline median runtime went from 45 minutes to 4 minutes.
- Pipeline median runtime
- 45m → 4m
- Concurrent builds supported
- 12 → 400+
- Time to first green build
- 8 weeks
- Beta cohort churn
- -60%
The problem
The company’s existing CI/CD product was a single-tenant Jenkins instance per customer, deployed manually by a solutions engineer. Median pipeline time was 45 minutes because every step ran in a serial VM with no layer cache. They’d lost two enterprise deals because the prospect’s security team wouldn’t approve Jenkins on their infrastructure.
Engineering leadership needed a multi-tenant platform that ran in the company’s own cloud, hit sub-five-minute pipeline times, and shipped within a quarter. They had three backend engineers who’d never shipped infrastructure-as-a-service before.
The work
I led the architecture and worked alongside their team for the duration. The plan was a Kubernetes-native runner pool with BuildKit for layer-cached builds, NATS for job distribution, and a Postgres-backed control plane. Each tenant got a namespace with quotas; jobs were dispatched via a queue and scheduled against the pool with priority-based preemption for paying customers.
The build cache was the hard part. We settled on a content-addressed blob store with per-tenant prefix isolation, populated lazily on first hit and evicted by LRU. The first integration test that ran against the new system had a 4-minute pipeline time — and the team thought we’d rigged it.
The outcome
The platform launched in beta with twelve design-partner customers. Median pipeline runtime held at 4 minutes across the cohort; one customer reported a 90% drop in CI compute cost after migrating. The company closed two of the enterprise deals that had previously been blocked on Jenkins approval, and the churn rate in the beta cohort was 60% lower than the legacy product’s launch cohort.
The team of three now runs the platform with one SRE. They hired a fourth in month five, not because they had to, but because the roadmap called for it.