Six-Month AKS Platform Assessment for Delta Data (now part of BetaNXT)

Six-Month AKS Platform Assessment for Delta Data (now part of BetaNXT)

Columbus, GA · Austin, TX

Executive Summary

The Client

Delta Data (Columbus, GA — founded 1985, now part of BetaNXT) builds the operational software behind mutual-fund processing. Its platforms have served 4 of the top 10 US banks, 3 of the top 5 retirement recordkeepers, and 23 of the top 25 US asset managers.

The Engagement

A six-month, five-workstream AKS platform and application assessment covering three revenue-critical applications — baseline the operational maturity, remediate quick wins in-flight, and hand the internal team a Target State Operating Model plus a phased roadmap they could execute without us.

47 findings
Baselined and prioritized across 5 workstreams — 11 Critical, 18 High
1/3 remediated
Quick Wins closed and verified during the engagement itself
3 apps
Revenue-critical mutual-fund applications audited end to end

Solution Implemented

Full platform baseline — production, QA, and POC AKS clusters plus three application audits across Observability, Logging, Cost/FinOps, Reliability, and Security. 47 findings captured: 11 Critical, 18 High, 13 Medium, 5 Low.

Quick Wins remediated in-flight (~1/3 of findings) — kube-prometheus-stack baseline with recording rules; centralized Azure Log Analytics with structured field extraction; OpenTelemetry instrumentation; Azure Key Vault CSI driver with legacy-secret migration; container-registry vulnerability-scan baseline; HPA ceilings and minReplicas aligned to SLOs; resource requests/limits; cluster-autoscaler hardening.

Cost levers identified — node utilization under 25% across user node pools; right-sizing and spot-scope strategy so eviction risk no longer equals availability risk; FinOps tagging policy and a monthly reporting cadence established with Finance.

Target State Operating Model + phased roadmap — 90-day, 6-month, and 12-month tracks with named owners, acceptance criteria, and ROI estimates: replicas ≥2 with PodDisruptionBudgets, default-deny NetworkPolicies, least-privilege RBAC, Kyverno/Gatekeeper admission control, SLO-based alerting.

Handover, not lock-in — runbook template adopted by the platform team, all audit source data in the client's own GitHub, deliverables in the client's SharePoint, and a plan to stand up their internal Platform Engineering function.

Outcomes Expected

47 findings baselined and prioritized — 11 Critical, 18 High, across five workstreams

~1/3 remediated before engagement close — each verified in its environment at sign-off

Observability stood up — kube-prometheus-stack + centralized logging where only a single APM vendor existed before

A roadmap the internal team executes without us — the no-lock-in model in practice

▸ Delta Data was subsequently acquired by BetaNXT (2025)

The Customer's Situation

Delta Data's Azure Kubernetes Service platform hosted three revenue-critical applications processing mutual-fund data for major US banks, recordkeepers, and asset managers. The platform supported production volume — but on a thin margin: single-replica production workloads, critical-path services on spot nodes, no PodDisruptionBudgets, no NetworkPolicies, cluster-admin service accounts, and observability limited to a single APM vendor. Leadership needed an honest baseline, a target operating model, and a roadmap their internal team could own.

The Six-Month Assessment

THNKBIG ran a phased, five-workstream assessment — Observability, Logging, Cost/FinOps, Reliability, and Security — across the production, QA, and POC clusters and all three applications. The engagement produced a maturity scorecard, per-application audit reports, a full cluster audit, an observability and logging review, a FinOps analysis, and a reliability and security review — 47 discrete findings, each classified and prioritized.

Fixed In-Flight, Not Just Reported

Roughly a third of the findings were remediated during the assessment itself: a kube-prometheus-stack observability baseline, centralized Azure Log Analytics with structured extraction, OpenTelemetry instrumentation, Azure Key Vault CSI with legacy-secret migration, a container-registry vulnerability-scan baseline, HPA and resource-limit alignment, cluster-autoscaler hardening, and a FinOps tagging policy with a monthly reporting cadence established with Finance. Every quick win was verified in its environment before sign-off.

A Roadmap Their Team Owns

The engagement closed with a Target State Operating Model and a 90-day / 6-month / 12-month roadmap — replicas ≥2 with PodDisruptionBudgets cluster-wide, default-deny NetworkPolicies, least-privilege RBAC, Kyverno/Gatekeeper admission control, SLO-based alerting, and a spot-instance strategy that keeps eviction risk out of the critical path. Each item carried a named owner, acceptance criteria, and an ROI estimate. All artifacts were handed over in Delta Data's own SharePoint and GitHub, with a runbook template their platform team adopted — no lock-in, by design. Delta Data was subsequently acquired by BetaNXT in 2025.

About THNKBIG

THNKBIG is an Austin, TX-based Platform Engineering consultancy specializing in Kubernetes, cloud-native, and production AI infrastructure — on-prem and cloud, vendor-neutral, with ownership transferred to your team. Specialties: platform assessments, GPU/on-prem AI, Kubernetes, software supply-chain security, and compliance readiness (FedRAMP, HIPAA, SOC 2, PCI, ATO).

Industry Context

Sector-Specific Challenges

Financial institutions operate under intense regulatory scrutiny while facing pressure to modernize legacy systems and deliver competitive digital experiences. These organizations must maintain 99.99% uptime for transaction processing, protect against sophisticated cyber threats, and ensure complete auditability of all system changes and data access.

Technical Considerations

Technical requirements for financial services infrastructure include real-time transaction processing with sub-millisecond latency, comprehensive data encryption at rest and in transit, multi-region disaster recovery with automatic failover, and immutable audit trails for regulatory examinations. Systems must support complex compliance reporting and risk management analytics.

Regulatory Environment

Financial services infrastructure typically must comply with SOC 2 Type II, PCI DSS for payment processing, GLBA for consumer data protection, and often additional requirements like FINRA rules for broker-dealers or OCC guidelines for banks.

Frequently Asked Questions

How do you approach client engagements?

Every engagement begins with a thorough discovery phase to understand your current state, business objectives, and constraints. We develop tailored recommendations rather than applying one-size-fits-all solutions. Our consultants work alongside your team to transfer knowledge and build sustainable capabilities. We measure success by business outcomes, not just technical deliverables.

What ROI can we expect from this type of engagement?

Organizations typically see significant improvements across multiple dimensions. Common outcomes include 50-80% reduction in deployment time, 30-50% decrease in infrastructure costs, 60-90% reduction in incident resolution time, and substantial improvements in developer productivity. The specific ROI depends on your starting point and investment level, which we help quantify during the assessment phase.

Related Solutions

This case study demonstrates our expertise in the following service areas. Learn more about how we can help your organization achieve similar results.

Cloud Complexity is a Problem — Until You Have the Right Team

From compliance automation to Kubernetes optimization, we help enterprises transform infrastructure into a competitive advantage.

Talk to a Cloud Expert