Back to insights

// INSIGHTS · ARCHITECTURE

Avoiding the distributed monolith

How not to spend years and a lot of money rebuilding the system you were trying to replace. Five early design decisions, and what to do about each.

Monoliths eventually hold organisations back, frustrating customers and employees alike. The legacy platform takes six weeks to release a two line change, teams block each other, and one slow reporting query can take the customer portal down. Microservices promise independent teams, independent releases and better failure isolation.

Three years and a sizeable budget later, many programmes have forty+ services, a Kubernetes cluster, a service mesh and an API gateway. They also have a six week release cycle, teams that still block each other and reporting queries that take customer portals down in new and harder to diagnose ways. The industry calls this a distributed monolith.

"All the coupling of the original, plus network latency and a larger cloud bill."

This is rarely a failure of architecture or engineering talent. The coupling comes from a few design decisions made early, often without enough rigour because of delivery pressure.

// DESIGN DECISION 01Decomposing services along technical lines

The easiest way to decompose a monolith is along its existing technical lines: a presentation tier, an orchestration tier, a validation component, a data access layer and a service per database entity. It’s quick, and teams keep working in technologies they know.

Business change doesn’t follow technical layers. Adding a new product type means changing five services, so one feature needs five teams and five backlogs. Conway’s law (systems mirror the communication structures of the organisations that build them) also works in reverse: a technically decomposed estate produces an organisation that needs a meeting to add a field.

Functional decomposition draws boundaries around business capabilities such as quoting, onboarding, billing and claims, so most changes land in one place. Domain driven design is the most common way to find them, but only if the model represents the whole business. Run the workshops with the people who know the current system best and you get the monolith’s modules under new names, coupling intact.

// DESIGN DECISION 02Retaining a shared operational data store

The most common decision is to keep the shared operational data store (ODS). The reporting team depends on it, so each new service gets its own codebase and pipeline and is pointed at the same schema.

The services are now independent in the way that flats sharing one boiler are independent. When the Orders team renames a column, Billing, Fulfilment and the nightly finance extract all have to test and agree the change.

A shared store is defensible during a transition. The problem comes when the transition state becomes the target architecture. Ask which service owns each table and who else writes to it. If the answer is "everyone", the monolith has simply moved into the database.

// DESIGN DECISION 03Services claim the business rules they need

Without a shared catalogue of rules, each service defines what it needs. Customer decides an active customer has logged in within 90 days; Orders that they have ordered this year; Billing that their account is paid up. Each is right in its own context, and the gap surfaces when the board pack and finance reports disagree by thousands of customers.

Domain driven design helps through bounded contexts (parts of the business in which a term has one precise meaning). Where definitions differ on purpose, document the translation. Where they differ by accident, name one owner and publish a governed contract. Duplicated logic is a fair price for independence.

// DESIGN DECISION 04Moving towards independent deployability

Continuous deployment feels like the natural next step, but the first three decisions make it almost impossible. When a change in one service only works alongside changes in several others, the programme creates a release train: every service deployed together after joint regression testing, overseen by a change advisory board.

That is a monolithic release process with added steps. Research from DORA (DevOps Research and Assessment) links small, frequent, independent deployments with better stability; a lockstep train gives most of that up.

Regulated sectors have legitimate reasons to gate change. But if a service can’t be deployed on a Tuesday afternoon without its neighbours, the coupling is architectural, and more release governance is unlikely to fix it. This is the kind of problem our fast flow work addresses.

// DESIGN DECISION 05Chaining synchronous calls

This decision is the most likely to cause customer-facing incidents. One customer request can pass from the gateway through the Orders API, the Customer API, the validation service and the data access service to the shared ODS.

If each of five services is available 99.9% of the time, the chain might only manage 99.5%: roughly three and a half hours of downtime a month against about 45 minutes for one service. Your old monolith at least had the decency to fail in one place.

The usual remedies are asynchronous messaging, events and local copies of the data each service needs. Eventual consistency takes getting used to, particularly for finance teams, but it beats availability set by the weakest links in a chain. Observability gives you the data to watch those boundaries and improve them.

// THE RESULTThe target architecture

Together, these decisions produce services cut along technical lines, each with its own version of the rules, all sharing one data store and one release train. The fix is to pace layer, decouple by domain and move ownership back to services: old concepts, but hard to apply at scale under pressure to deliver.

// BEFORE YOU BEGINAre microservices the right pattern?

Ask that question first. Microservices are not the only way to treat a monolith that is causing your organisation pain.

// TALK TO US

We’re supporting clients across several sectors with exactly these problems, some partway through a migration and some still deciding how to decompose their estate. If this sounds familiar, we’d be happy to share what we’ve learned.

This article was first published by Adam Cockburn on LinkedIn.

// LET'S TALK

Living with a distributed monolith?

Whether you are partway through a migration or still deciding how to decompose your estate, we can share what we have learned on similar programmes.