Platform Security Assessment & Infrastructure Control Plane for an Enterprise Cloud Communications Provider
Austin, TX
Executive Summary
The Client
An enterprise cloud communications provider running production infrastructure across multiple customer-managed data centers, with a platform team responsible for datacenter automation — DCIM, IPAM, storage, compute, and bare-metal provisioning — underpinning customer-facing services.
The Engagement
THNKBIG was brought in to assess the security and supply-chain posture of the platform team's datacenter-automation stack, then to design and build a modern, auditable control plane to replace brittle, ad-hoc tooling — with knowledge transfer so the customer's own team owns and operates it.
Solution Implemented
✔ Security assessment of a Python/FastAPI infrastructure API spanning 63 endpoints across DCIM, IPAM, storage, compute, and bare-metal — delivering a 75-finding remediation roadmap and a six-week sprint plan adopted by the platform team.
✔ Go, plugin-driven control plane with 17 provider integrations (Infoblox, Kubernetes, Cisco UCS/Redfish, DCIM/CMDB, Grafana), a 51-field CMDB with an 11-state asset lifecycle, Dex OIDC RBAC, and multi-datacenter federation.
✔ Bare-Metal-as-a-Service provisioning integrating GitLab CI, Tinkerbell, Cisco UCS service profiles, and Pure Storage — with HMAC-signed callbacks, idempotency guarantees, and a sub-30-minute p95 provisioning SLA.
✔ Software supply-chain hardening evaluated against 16 OpenSSF standards (SLSA, SBOM, Sigstore/cosign, Scorecard, secrets scanning), with a hardening roadmap, a 620-test quality-engineering plan, and a 12-factor modernization assessment.
✔ GitOps delivery (GitLab CI, Kaniko, Harbor, ArgoCD) with distroless multi-stage builds, Vault-agent secret injection, CloudNativePG PostgreSQL HA, and Prometheus/Grafana observability.
✔ Kubernetes consolidation onto bare-metal Rancher RKE2 and Harvester across multiple production data centers, resolving secrets-management, service-account, and workload-migration blockers.
Outcomes Expected
▸ 75-finding security remediation roadmap and six-week sprint plan adopted by the customer's platform team
▸ Go control plane with 17 provider integrations replacing brittle, ad-hoc automation with an auditable, RBAC-governed system
▸ Sub-30-minute p95 bare-metal provisioning via a fully automated BMaaS pipeline
▸ Supply-chain posture measured against 16 OpenSSF standards with a prioritized hardening roadmap
▸ Operational ownership transferred to the customer's team — runbooks, documentation, and GitOps workflows all customer-owned
The Customer's Situation
The platform team had grown a datacenter-automation stack organically: a large Python/FastAPI API surface (63 endpoints across DCIM, IPAM, storage, compute, and bare-metal), scripts, and manual runbooks. It worked, but no one could fully attest to its security posture, its supply-chain integrity, or how it would behave under change. Provisioning a new bare-metal server was slow and inconsistent, and the team wanted an auditable control plane they could own — not another black box.
Assess First, Then Build
THNKBIG started with a fixed-scope assessment rather than a rewrite. The existing API was audited endpoint by endpoint, producing a 75-finding remediation roadmap and a six-week sprint plan the platform team adopted directly. In parallel, the software supply chain was measured against 16 OpenSSF standards (SLSA, SBOM, Sigstore/cosign, Scorecard, secrets scanning) to establish an objective baseline and a prioritized path forward.
An Auditable, Owned Control Plane
The new control plane was built in Go as a plugin-driven system with 17 provider integrations (Infoblox, Kubernetes, Cisco UCS/Redfish, DCIM/CMDB, Grafana), a 51-field CMDB with an 11-state asset lifecycle, Dex OIDC RBAC, and multi-datacenter federation. Bare-metal provisioning was automated end to end with GitLab CI, Tinkerbell, Cisco UCS service profiles, and Pure Storage — HMAC-signed callbacks and idempotency guarantees brought p95 provisioning under 30 minutes. Delivery ran on GitOps (GitLab CI, Kaniko, Harbor, ArgoCD) with distroless builds, Vault-agent secret injection, CloudNativePG HA, and Prometheus/Grafana observability.
Engagement Results
The platform team adopted the 75-finding remediation roadmap; the Go control plane replaced ad-hoc automation with an auditable, RBAC-governed system; bare-metal provisioning reached sub-30-minute p95; and Kubernetes across multiple data centers was consolidated onto bare-metal Rancher RKE2 and Harvester. Ownership — runbooks, documentation, and GitOps workflows — was transferred to the customer's team. Detailed findings, architecture, and metrics are available to qualified buyers under NDA.
About THNKBIG
THNKBIG is an Austin, TX-based Platform Engineering consultancy specializing in Kubernetes, cloud-native, and production AI infrastructure. We build around how your business actually runs and transfer ownership to your team. Specialties: platform engineering, on-prem and cloud Kubernetes, software supply-chain security, infrastructure automation, and compliance readiness (FedRAMP, HIPAA, SOC 2, PCI, ATO).
Industry Context
Sector-Specific Challenges
Technology companies must maintain development velocity while ensuring platform reliability and security. These organizations face challenges in scaling microservices architectures, managing complex deployment pipelines, and maintaining developer productivity without compromising production stability.
Technical Considerations
Technology infrastructure requires support for continuous deployment practices, comprehensive observability across distributed systems, feature flag management, and chaos engineering capabilities. Systems must enable rapid experimentation while maintaining service level objectives.
Regulatory Environment
Technology companies typically must demonstrate SOC 2 Type II compliance for enterprise customers, GDPR compliance for European users, and often industry-specific certifications based on customer requirements.
Frequently Asked Questions
How do you approach client engagements?
Every engagement begins with a thorough discovery phase to understand your current state, business objectives, and constraints. We develop tailored recommendations rather than applying one-size-fits-all solutions. Our consultants work alongside your team to transfer knowledge and build sustainable capabilities. We measure success by business outcomes, not just technical deliverables.
What ROI can we expect from this type of engagement?
Organizations typically see significant improvements across multiple dimensions. Common outcomes include 50-80% reduction in deployment time, 30-50% decrease in infrastructure costs, 60-90% reduction in incident resolution time, and substantial improvements in developer productivity. The specific ROI depends on your starting point and investment level, which we help quantify during the assessment phase.
Related Solutions
This case study demonstrates our expertise in the following service areas. Learn more about how we can help your organization achieve similar results.
Cloud Complexity is a Problem —
Until You Have the Right Team
From compliance automation to Kubernetes optimization, we help enterprises transform infrastructure into a competitive advantage.
Talk to a Cloud Expert