Drezen Technology

Helios Analytics

Scaling a B2B SaaS platform through 8x growth without a rewrite

Targeted performance and isolation work absorbed 8x tenant growth, cut p95 latency 78%, and took deployment from fortnightly to daily.
SaaS & Technology6 months5 people2025

Results

Reduction in p95 API latency
78%Reduction in p95 API latency
Deployment frequency
14 days → 1 dayDeployment frequency
Tenant growth absorbed
8xTenant growth absorbed

The challenge

A B2B analytics SaaS had found product-market fit and was onboarding enterprise customers considerably larger than anything the platform had been designed for. A single large tenant running a heavy report could degrade response times for every other customer, and the team was fielding performance complaints weekly.

The internal assumption was that a rewrite onto a microservices architecture was needed. The engineering team had begun planning it, which would have consumed most of the following year at exactly the moment the commercial team needed feature velocity.

What we did

  1. 01

    Measured before rebuilding

    Three weeks of query-level profiling and distributed tracing showed that eleven query patterns accounted for 84% of database load, and that two tenants generated most of the contention. The rewrite was not the required fix.

  2. 02

    Fixed the queries and the indexes

    Targeted index work, query rewrites and pre-aggregated materialised views for the heaviest report patterns — a fraction of the effort of a rewrite, with most of the latency improvement.

  3. 03

    Isolated tenant workloads

    Heavy analytical work moved to a separate connection pool and worker tier with per-tenant concurrency limits, so no single customer can starve the others regardless of what they run.

  4. 04

    Matured the delivery pipeline

    Build caching, test parallelisation, preview environments per pull request and canary deployment with automated rollback took the release cycle from fortnightly and nervous to daily and routine.

  5. 05

    Established SLOs and error budgets

    Per-service objectives agreed with the commercial team, with burn-rate alerting replacing an alert set that had trained everyone to ignore it.

The outcome

p95 API latency dropped 78% and the platform absorbed 8x tenant growth over the following year on substantially the same architecture. The planned rewrite was cancelled, releasing roughly a year of engineering capacity back to product work.

Deployment frequency moved from fortnightly to daily, and change failure rate fell as batch sizes shrank. The engineering team now runs the SLO review cadence themselves.

Services

How we did it

All services

SaaS Product Development

Multi-tenant platforms with billing, onboarding and analytics built in.

Learn more

DevOps & CI/CD Automation

Pipelines that make deploying boring — and therefore frequent.

Learn more

Site Reliability Engineering (SRE)

SLOs, observability and incident practice that make uptime predictable.

Learn more

More

Other case studies

View all
Fintech & Financial Services9 months

Rebuilding a lending origination platform around an event-sourced ledger

An event-sourced origination platform replacing a spreadsheet-and-email process, with decisioning, audit trail and reconciliation built in.

Median time to credit decision
4.2 days → 6 hrsMedian time to credit decision
Applications processed per underwriter
3.1xApplications processed per underwriter
Decisions reconstructable for audit
100%Decisions reconstructable for audit
Read the case study
Fintech & Financial Services11 months

Migrating a regional bank to AWS without a maintenance window

43 workloads moved from two ageing data centres to AWS in eleven months, with a governed landing zone and a 34% run-rate reduction.

Unplanned outages during migration
0Unplanned outages during migration
Infrastructure run-rate reduction
34%Infrastructure run-rate reduction
Environment provisioning time
11 days → 40 minEnvironment provisioning time
Read the case study
Logistics & Supply Chain7 months

A supply chain control tower that surfaces exceptions before they land

A real-time visibility platform unifying 19 carrier feeds, with predictive ETAs and exception routing that reaches an operator while options still exist.

Reduction in late-delivery penalties
61%Reduction in late-delivery penalties
Systems an operator monitors
19 → 1Systems an operator monitors
Average earlier exception detection
4.5 hrsAverage earlier exception detection
Read the case study

Tell us what you are trying to build

A 30-minute call is usually enough for both of us to know whether this is a fit. If it is not, we will point you somewhere better.

sales@drezentechnology.comUsually replies within one business day