Case study

Self-Service Kubernetes Platform for Financial Services

A major financial institution needed a self-service Kubernetes platform providing hosting, monitoring and scaling for multiple development teams.

Industry
Financial Services
Duration
18 months
Self-service Kubernetes platform and its move from VPC peering to hub and spokeThe upper half contrasts two network topologies. On the left, five VPCs are peered to one another in a full mesh, so each additional team adds a connection to every existing VPC. On the right, the same VPCs each hold a single link to a central transit gateway acting as a hub. The lower half shows how teams use the result: development teams reach a self-service platform that provisions EKS clusters as spoke VPCs behind that transit gateway. Beneath them, the platform applies Vault for secrets, Falco for runtime container security and Istio as a service mesh to every cluster, while OpenTelemetry collects traces and metrics from every cluster and feeds Grafana dashboards and alerts.Before: VPC peeringVPCVPCVPCVPCVPCEach new team adds a linkto every other VPCAfter: hub and spokeVPCVPCVPCVPCTransit GatewayOne connection per team,routed through the hubDevelopmentteamsSelf-serviceplatformAutomated clustersTransit GatewayNetwork hubSpoke VPCsEKS clusterEKS clusterEKS clusterApplied by the platform to every clusterVaultSecrets managementFalcoRuntime container securityIstioService meshOpenTelemetryTraces and metricsGrafanaDashboards and alertsTelemetry from every cluster
A simplified view of the network change and what teams self-served onto it. The peering mesh is drawn with five VPCs to show the shape of the problem, not the count the organisation ran. Illustrative, not as-built documentation.

The challenge

The organisational problem was the real one. A central operations team had become a bottleneck: every deployment, every scaling decision and every new environment went through the same small group. Teams waited days for changes that should take minutes, and operations spent its time on repetitive requests instead of on the platform itself.

The approach

We built an automated cluster platform that fused software development with operations, so that the safe path for a development team was also the fast one.

The networking was restructured from VPC peering to a hub-and-spoke model with transit gateways. Peering does not scale organisationally: each new team compounds the connection count and the review burden.

Security controls were integrated, not bolted on. Vault handled secrets and Falco handled runtime container security, with policy enforced by the platform instead of by review. Observability used OpenTelemetry and Grafana and was in place from the start, because a self-service platform teams cannot see into only moves the support burden somewhere else.

The outcome

  • Deployment time for development teams fell from days to minutes
  • Platform uptime reached 99.9% through automated scaling and monitoring
  • Self-service capability reduced operational overhead by around 60%
  • Security posture improved through integrated scanning and policy enforcement

The 60% reduction in operational overhead is the number worth dwelling on: it represents a central team no longer spending most of its week on requests that the platform could answer itself.

Next step

Facing a problem like this one?

We work on platform engagements where the constraints are real and the outcome is measurable. Describe yours and we will tell you whether we are the right people for it.