Executive Summary
The Client
An enterprise standing up its first production on-prem AI platform — GPU hardware already in the rack, and a hard deadline to get real workloads running for its own customers.
The Engagement
A six-week rapid-strike build: THNKBIG stood up a production-ready, CIS-hardened NVIDIA GPU AI cluster on RKE2 and worked alongside the client's application owners to ship an internal AI application on top of it. This contrasts with a longer, procurement-gated engagement (where hardware is still being acquired and the work is advisory over many months) — here the mandate was speed to a hardened, production-ready platform.
Solution Implemented
✔ RKE2 Kubernetes cluster with dedicated GPU node groups, provisioned and production-ready.
✔ NVIDIA GPU Operator — driver lifecycle, device plugin, and GPU scheduling on bare-metal nodes.
✔ Storage — Longhorn on Pure Storage for resilient, performant persistent volumes.
✔ Private registry — Harbor for signed, scanned images inside the perimeter.
✔ Observability — Prometheus, Grafana, and NVIDIA DCGM for GPU-level telemetry.
✔ Hardening — kube-bench + CIS Benchmarks remediated to pass before go-live.
✔ Load testing — k6 to validate the platform under production-like traffic.
✔ Sovereign model hosting — self-hosted open-weight models (Qwen, MiniMax, Gemma, Kimi) via Hugging Face and NVIDIA NeMo, all inference inside the client perimeter with no third-party model egress.
✔ Application delivery — worked directly with the client's app owners to build and ship an internal AI application serving their customers.
Outcomes Expected
▸ Production-ready on-prem GPU AI cluster in six weeks — RKE2 + NVIDIA GPU Operator + GPU node groups
▸ CIS-hardened before go-live — kube-bench + CIS Benchmarks remediated; k6 load-tested
▸ Sovereign inference — self-hosted open models via Hugging Face / NeMo, zero third-party model egress
▸ An internal AI application shipped with the client's app owners, serving their customers
▸ Owned by the client's team at handoff — runbooks and documentation included
The Customer's Situation
The GPU hardware was already racked, but a bare server is not an AI platform. The client needed production-ready, hardened Kubernetes infrastructure that could take real AI workloads — fast — and an internal AI application on top of it serving their own customers. The mandate was speed without cutting the corners that matter: security, storage resilience, and GPU utilization.
The Six-Week Rapid Strike
THNKBIG built an RKE2 cluster with dedicated GPU node groups, layered on the NVIDIA GPU Operator for driver lifecycle and GPU scheduling, and wired Longhorn on Pure Storage for resilient persistent volumes. Harbor provided a private, scanned image registry inside the perimeter; Prometheus, Grafana, and NVIDIA DCGM gave GPU-level observability from day one.
Hardened and Load-Tested Before Go-Live
Security was designed in, not bolted on: the cluster was assessed with kube-bench against the CIS Benchmarks and remediated to pass before it took production traffic. k6 load tests validated the platform under production-like demand so go-live was a non-event.
Sovereign Models and a Shipped Application
Inference ran on self-hosted open-weight models — Qwen, MiniMax, Gemma, and Kimi — served via Hugging Face tooling and NVIDIA NeMo, entirely inside the client's perimeter with no third-party model egress. THNKBIG then worked hand-in-hand with the client's application owners to build and ship an internal AI application on the new platform, serving their customers.
Engagement Results
In six weeks the client went from racked hardware to a production-ready, CIS-hardened, load-tested on-prem GPU AI cluster running sovereign inference — with an internal AI application live and owned by their own team. Detailed sizing, benchmark results, and utilization metrics are available to qualified buyers under NDA.
About THNKBIG
THNKBIG is an Austin, TX-based Platform Engineering consultancy specializing in Kubernetes, cloud-native, and production AI infrastructure — on-prem and cloud, vendor-neutral, with ownership transferred to your team. Specialties: GPU/on-prem AI, sovereign inference, Kubernetes, software supply-chain security, and compliance readiness (FedRAMP, HIPAA, SOC 2, PCI, ATO).
Industry Context
Sector-Specific Challenges
Technology companies must maintain development velocity while ensuring platform reliability and security. These organizations face challenges in scaling microservices architectures, managing complex deployment pipelines, and maintaining developer productivity without compromising production stability.
Technical Considerations
Technology infrastructure requires support for continuous deployment practices, comprehensive observability across distributed systems, feature flag management, and chaos engineering capabilities. Systems must enable rapid experimentation while maintaining service level objectives.
Regulatory Environment
Technology companies typically must demonstrate SOC 2 Type II compliance for enterprise customers, GDPR compliance for European users, and often industry-specific certifications based on customer requirements.
Frequently Asked Questions
How do you approach client engagements?
Every engagement begins with a thorough discovery phase to understand your current state, business objectives, and constraints. We develop tailored recommendations rather than applying one-size-fits-all solutions. Our consultants work alongside your team to transfer knowledge and build sustainable capabilities. We measure success by business outcomes, not just technical deliverables.
What ROI can we expect from this type of engagement?
Organizations typically see significant improvements across multiple dimensions. Common outcomes include 50-80% reduction in deployment time, 30-50% decrease in infrastructure costs, 60-90% reduction in incident resolution time, and substantial improvements in developer productivity. The specific ROI depends on your starting point and investment level, which we help quantify during the assessment phase.
Related Solutions
This case study demonstrates our expertise in the following service areas. Learn more about how we can help your organization achieve similar results.
Cloud Complexity is a Problem —
Until You Have the Right Team
From compliance automation to Kubernetes optimization, we help enterprises transform infrastructure into a competitive advantage.
Talk to a Cloud Expert