Skip to content
KAZES.studio — home
All projects

Developer infrastructure

Concept study

Making deploys boring across forty repositories

A single release path for a growing engineering organisation, where the previous approach made every deploy a coordinated event involving three teams.

Engagement
Embedded engineer
Duration
Ongoing, eight months to first restructure
Year
2025
Client
Not applicable
010203040506
Data path, end to end.. This is a generated schematic, not a product screenshot.

The challenge

What made this difficult.

Services had accumulated their own pipelines, their own deployment triggers, and their own rollback procedures. A change touching a shared library required a coordinated release across teams, and the cost of that coordination had made teams batch unrelated changes into larger, riskier deploys. The bottleneck was not code review — it was the release path itself.

Constraints

Non-negotiables we designed around.

01

No big-bang migration

Incremental migration mattered more than a clean end state. Forty services could not be frozen while the platform changed.

02

Different runtimes already in production

The constraint applied to containerised services and to scheduled batch jobs with entirely different failure characteristics.

03

Existing ownership boundaries

Team boundaries had to keep matching on-call responsibility. A pipeline change that obscured ownership would be rejected.

04

Audit requirements on release events

Every production change needed an attributable record, and that record had to survive platform changes.

Approach

How we would build it.

Started with one service, end to end

Rather than designing the target platform on paper, we migrated a single representative service completely — build, test, canary, promote, rollback — and used the friction as the requirements document for everything else.

Made provenance a build property

Every artefact carries the source revision, the build inputs, and the change record with it. Promotion becomes a reference to a verified artefact rather than a rebuild, which is what removed the coordinated-release requirement.

Progressive rollout as the default, not an option

Canary analysis runs on every production change with automatic halt thresholds. Small teams stopped needing a manual go/no-go decision they had no context to make well.

Rollback designed before rollout

Because rollout is a shift of traffic between immutable artefacts, rollback is a traffic operation. We rehearsed it under failure conditions rather than assuming it worked.

Migrated on an agreed service queue

Teams opted in as they had capacity. A thin compatibility layer meant both patterns ran in parallel, so the migration never became a coordination exercise.

Architecture

How the pieces fit together.

Architecture

Unified release path. Promotion shifts traffic between immutable artefacts, so rollback is a traffic operation.

  1. Source

    • Monorepo
    • Service modules
    • Shared libraries
    • Infrastructure definitions
  2. Build

    • Hermetic builds
    • Content-addressed artefacts
    • SBOM generation
    • Vulnerability scanning
  3. Verify

    • Unit and integration
    • Contract tests
    • Environment gates
    • Artefact attestation
  4. Release

    • Progressive rollout
    • Automatic halt thresholds
    • Traffic shifting
    • Instant rollback
  5. Operate

    • Unified telemetry
    • Change correlation
    • Ownership metadata
    • Incident annotations

Stack

What it would run on.

Language

  • Go
  • TypeScript
  • Python
  • Bash

Platform

  • Container images
  • Declarative infrastructure
  • OCI registries
  • Secret management

Release

  • Progressive delivery controller
  • Service mesh traffic management
  • Policy-as-code
  • CI with attestation

Expected outcomes

What success would look like.

Coordinated releases

Eliminated

Design goal: shared-library changes no longer require a multi-team release window.

Rollback

Traffic shift

Because promotion references immutable artefacts rather than rebuilding, rollback no longer depends on a reversible build.

Adoption

Opt-in queue

Both release patterns ran in parallel throughout, so migration was never a freeze-window negotiation.

What we learned

The conclusions we would carry forward.

Migrating one service completely was worth more than designing the whole platform — the failure modes showed up immediately.

Making builds hermetic and artefacts immutable is what turned rollback from a build problem into a routing problem.

An agreed migration queue kept platform work from turning into a coordination tax on every team.