On-Prem K8s + AI Cluster Buildout for a Top-5 Cloud Contact Center SaaS Vendor
Austin, TX
Executive Summary
The Client
A top-5 cloud contact center SaaS vendor serving thousands of enterprise customers from customer-managed data centers. The customer's platform team had strong general platform expertise but needed named senior engineers to deliver a parallel two-cluster on-prem Kubernetes topology — one for the production platform tier, one for AI/ML inference powering customer-facing AI features.
The Engagement
Senior Platform Engineers, US-based, reporting directly to the customer's VP of Infrastructure. Two on-prem Kubernetes clusters delivered in parallel over a 12-month engagement with zero customer-visible incidents during build-out.
Solution Implemented
✔ Production Kubernetes cluster on bare metal — full lifecycle management, customer-facing workloads.
✔ AI/ML Kubernetes cluster running customer-facing AI features (real-time transcription, agent assist, post-call summarization) inside the same data center perimeter, no third-party model-provider egress.
✔ Network policy + multi-tenancy hardening — trust-boundary separation without crossing trust boundaries.
✔ Observability and SLO design layered over the customer's existing APM rather than replacing it.
✔ CI/CD + GitOps so the customer's platform team could ship without operator intervention after handoff.
✔ Runbooks, training, and operational documentation for full ownership at handoff.
Outcomes Expected
▸ Two-cluster topology delivered in parallel: production platform cluster + AI/ML cluster, both running on customer bare metal
▸ Customer-facing AI features shipped to production traffic — real-time transcription, agent assist, post-call summarization
▸ Sovereignty posture maintained: all inference on customer bare metal, no third-party model-provider egress
▸ Operational handoff: customer's platform team fully owns both clusters at handoff
▸ 12-month senior engagement — engagement value available under NDA
The Customer's Situation
A multi-tenant cloud contact center platform serving thousands of enterprise customers needed AI features (real-time transcription, agent assist, post-call summarization) shipped from prototype to customer-facing traffic. The architecture needed to support both the core platform and the AI workload inside the customer's own data centers, with strict data sovereignty requirements that no hyperscaler-managed solution could satisfy.
Three constraints shaped the architecture.
1. On-prem, in customer-managed data centers.
Bare metal Kubernetes, not AWS, not Azure — with change-control windows the upstream Kubernetes minor-version cadence had to be engineered around. Clusters were deliberately pinned to N-1 for several months and upgraded as paired windows opened.
2. Two clusters in parallel.
One production platform cluster, one AI/ML cluster. Different upgrade cadences, different network and trust boundaries, different compliance posture — and a zero-customer-data-leave-the-perimeter requirement applied to both.
3. Senior, single-thread delivery.
The customer platform team had strong general platform expertise but needed named senior engineers accountable for the on-prem K8s + AI cluster topology specifically. THNKBIG ran the engagement reporting directly to the customer's VP of Infrastructure, with no SI hand-holding.
Why This Is a Data Sovereignty + AI Infrastructure Engagement
AI inference at production scale does not require hyperscaler egress. The customer chose on-prem K8s + AI infrastructure over a managed-model-provider path specifically to maintain data sovereignty. AI inference happened inside their perimeter, on bare metal, on Kubernetes, with the customer's compliance posture unchanged.
That's the design pattern more regulated enterprise buyers are asking for in 2026: AI capabilities without surrendering data custody. The same architecture — sovereign inference, network-isolated AI cluster, no model-provider data egress — applies to financial services, healthcare, federal, and defense use cases that are explicitly blocked from sending customer data to third-party model APIs.
Defended Architectural Tradeoffs
Service mesh evaluated and not adopted in production. Network policy plus multi-tenancy hardening delivered the trust-boundary separation the customer needed without the operational tax of a sidecar mesh at the workload counts in scope. The mesh remains evaluated for re-introduction if east-west traffic grows.
Observability deliberately layered, not unified. The customer's platform team retained ownership of their existing APM layer; THNKBIG built the cluster- and workload-level SLO instrumentation on top rather than displacing the incumbent tooling.
Engagement Results
Engagement length: 12 months from kickoff to handoff, with advisory engagement ongoing. Workload classes shipped: production platform cluster; AI/ML cluster powering customer-facing AI features. Sovereignty posture: all inference on customer bare metal, no third-party model-provider egress, full data custody retained by customer throughout. Operational handoff: customer's platform team owned both clusters at handoff; runbooks, on-call rotation, and upgrade procedures all customer-owned. Reporting line: THNKBIG reported directly to the customer's VP of Infrastructure for the engagement duration.
Detailed cluster sizing, deployment cadence, SLO adherence, GPU utilization, and incident metrics are available to qualified buyers under NDA.
Frequently Asked Questions
How long does it take to stand up a production on-prem Kubernetes cluster?
For a single production cluster with full GitOps, observability, and handoff: typically 3-6 months from kickoff to production traffic, depending on data-center complexity, network segmentation, and customer-platform-team involvement. A two-cluster engagement with build-out plus multi-month handoff typically runs 12 months.
How much does a 12-month Kubernetes consulting engagement cost?
Engagement cost depends on team size and duration; a typical 12-month engagement of two to three senior engineers is a multi-hundred-thousand-dollar investment. Specific rates and total contract value are shared under NDA. Every hour billed is a senior hour — we don't carry a junior bench.
What changes when Kubernetes runs on-prem instead of in the cloud?
Four things change operationally: data-center change-control windows become the de facto deployment cadence; network policy becomes a first-class concern because egress is no longer a default-allow hyperscaler posture; the GPU + AI serving stack runs on bare metal with tighter failure modes than cloud-grade autoscaling; the customer retains full data sovereignty, which is the entire reason on-prem is being chosen.
Can your engineers work alongside our existing platform team?
Yes. We embed at the engineer level — no project manager between your team and ours. We take on scoped problems (a new cluster topology, a cost reduction initiative, a compliance audit) and work in the customer's existing tools and workflows. The customer's platform team retained full ownership of both clusters at handoff.
Book a 30-minute scoping call at https://thnkbig.com/scope. We'll tell you in the first 15 minutes whether we're the right fit, and if we aren't, who is.
About THNKBIG
THNKBIG is an Austin, TX-based Platform Engineering consultancy specializing in Kubernetes and AI infrastructure. Every engineer on our engagements has 10+ years of production Kubernetes experience and is US-based. Specialties: on-prem Kubernetes, sovereign AI inference, multi-cluster governance, FinOps for GPU + token spend, data sovereignty architecture, and compliance readiness (FedRAMP, HIPAA, SOC 2, CMMC).
Frequently Asked Questions
How do you approach client engagements?
Every engagement begins with a thorough discovery phase to understand your current state, business objectives, and constraints. We develop tailored recommendations rather than applying one-size-fits-all solutions. Our consultants work alongside your team to transfer knowledge and build sustainable capabilities. We measure success by business outcomes, not just technical deliverables.
What ROI can we expect from this type of engagement?
Organizations typically see significant improvements across multiple dimensions. Common outcomes include 50-80% reduction in deployment time, 30-50% decrease in infrastructure costs, 60-90% reduction in incident resolution time, and substantial improvements in developer productivity. The specific ROI depends on your starting point and investment level, which we help quantify during the assessment phase.
Related Solutions
This case study demonstrates our expertise in the following service areas. Learn more about how we can help your organization achieve similar results.
Cloud Complexity is a Problem —
Until You Have the Right Team
From compliance automation to Kubernetes optimization, we help enterprises transform infrastructure into a competitive advantage.
Talk to a Cloud Expert