From a Bespoke Stack to Commodity Chef

We replaced a custom, home-grown configuration system on ~30,000 servers with commodity Chef — zero downtime — and the hardest part wasn't the engineering. It was the politics.

Simpler, clearer, supportable by anyone — and one ground truth that finally brought Dev, Infrastructure, Security, and Audit into agreement.

See How

The Problem

A custom stack almost nobody could understand — hiding a security blind spot.

A global payments platform ran its entire ~30,000-server bare-metal fleet on a home-grown configuration system: a package manager implemented entirely in userspace, in Puppet, built that way for one reason — the team wasn't allowed to run the agent as root. Fully 80–90% of that code existed only to work around the no-root rule.

Worse, it was a security blind spot. Audit and security controls read the installed operating-system packages — but all the real work happened in unmonitored userspace, behind a custom control plane almost nobody understood. The controls were effectively lying about what the machines were doing. And the premise didn't even hold: the hosts were single-tenant, so the isolation concern behind the no-root rule didn't apply. The company had run an 18-month search just to find people who could maintain the thing.

The Real Work Was Political

The technical fix was obvious — run standard Chef, as root, with real OS package management. Getting permission to do it, in a large and control-heavy organization, was the actual challenge.

So we won the change before touching the code. We went to Security first and showed them their own monitoring was blind to the userspace where the real work happened — and that everyone would get what they actually wanted if the platform could just run Chef as root. Security agreed. Then they said we'd never get it past Audit.

We gave Audit the same case: what you're tracking is bogus, because the real work isn't where you're looking. Audit agreed, too — and said we'd never get it past Security. With both already bought in, each certain the other would block it, the objection evaporated. There was nothing left to say no.

The Migration

With the politics settled, the engineering was the straightforward part. We re-engineered the fleet from the userspace-Puppet hack to a standard Chef deployment — agent running as root, real operating-system package management — and migrated roughly 30,000 bare-metal servers with zero downtime, on a live, PCI DSS compliant payments platform.

The Result

We traded the appearance of safety for something real. The monitoring blind spot closed, and Dev, Infrastructure, Security, and Audit finally stood on one standard system that produced truthful data.

30,000 Servers, Zero Downtime

Live PCI DSS payments platform

The whole fleet moved from bespoke userspace Puppet to commodity Chef without an outage.

Supportable by Anyone

Not just the two people who built it

Any engineer they hired — or anyone who simply read the code — could now run it, instead of a custom control plane only a couple of people understood.

One Ground Truth

Dev, Infra, Security, Audit aligned

Four organizations that had been talking past each other finally reasoned about the same system, seeing the same real data.

The lesson we carry from it: the win wasn't the tooling — it was making things simpler, clearer, and easier to support, and bringing separate teams onto a single ground truth. Do that, and everyone's job gets easier. The hard part is rarely the migration. It's the persuasion.

Have a bespoke stack nobody wants to touch?

We simplify complex systems down to well-understood, commodity foundations — and bring the teams around them into agreement.

Get in Touch