Platform Security Assessment & Infrastructure Control Plane for an Enterprise Cloud Communications Provider

Platform Security Assessment & Infrastructure Control Plane for an Enterprise Cloud Communications Provider

Austin, TX

Executive Summary

The Client

An enterprise cloud communications provider running production infrastructure across multiple customer-managed data centers, with a platform team responsible for datacenter automation — DCIM, IPAM, storage, compute, and bare-metal provisioning — underpinning customer-facing services.

The Engagement

THNKBIG was brought in to assess the security and supply-chain posture of the platform team's datacenter-automation stack, then to design and build a modern, auditable control plane to replace brittle, ad-hoc tooling — with knowledge transfer so the customer's own team owns and operates it.

75 findings
Security remediation roadmap adopted by the platform team
17 integrations
Plugin-driven infrastructure control plane
<30 min p95
Bare-metal server provisioning SLA

Solution Implemented

Security assessment of a Python/FastAPI infrastructure API spanning 63 endpoints across DCIM, IPAM, storage, compute, and bare-metal — delivering a 75-finding remediation roadmap and a six-week sprint plan adopted by the platform team.

Go, plugin-driven control plane with 17 provider integrations (Infoblox, Kubernetes, Cisco UCS/Redfish, DCIM/CMDB, Grafana), a 51-field CMDB with an 11-state asset lifecycle, Dex OIDC RBAC, and multi-datacenter federation.

Bare-Metal-as-a-Service provisioning integrating GitLab CI, Tinkerbell, Cisco UCS service profiles, and Pure Storage — with HMAC-signed callbacks, idempotency guarantees, and a sub-30-minute p95 provisioning SLA.

Software supply-chain hardening evaluated against 16 OpenSSF standards (SLSA, SBOM, Sigstore/cosign, Scorecard, secrets scanning), with a hardening roadmap, a 620-test quality-engineering plan, and a 12-factor modernization assessment.

GitOps delivery (GitLab CI, Kaniko, Harbor, ArgoCD) with distroless multi-stage builds, Vault-agent secret injection, CloudNativePG PostgreSQL HA, and Prometheus/Grafana observability.

Kubernetes consolidation onto bare-metal Rancher RKE2 and Harvester across multiple production data centers, resolving secrets-management, service-account, and workload-migration blockers.

Outcomes Expected

75-finding security remediation roadmap and six-week sprint plan adopted by the customer's platform team

Go control plane with 17 provider integrations replacing brittle, ad-hoc automation with an auditable, RBAC-governed system

Sub-30-minute p95 bare-metal provisioning via a fully automated BMaaS pipeline

Supply-chain posture measured against 16 OpenSSF standards with a prioritized hardening roadmap

Operational ownership transferred to the customer's team — runbooks, documentation, and GitOps workflows all customer-owned

The Customer's Situation

The platform team had grown a datacenter-automation stack organically: a large Python/FastAPI API surface (63 endpoints across DCIM, IPAM, storage, compute, and bare-metal), scripts, and manual runbooks. It worked, but no one could fully attest to its security posture, its supply-chain integrity, or how it would behave under change. Provisioning a new bare-metal server was slow and inconsistent, and the team wanted an auditable control plane they could own — not another black box.

Assess First, Then Build

THNKBIG started with a fixed-scope assessment rather than a rewrite. The existing API was audited endpoint by endpoint, producing a 75-finding remediation roadmap and a six-week sprint plan the platform team adopted directly. In parallel, the software supply chain was measured against 16 OpenSSF standards (SLSA, SBOM, Sigstore/cosign, Scorecard, secrets scanning) to establish an objective baseline and a prioritized path forward.

An Auditable, Owned Control Plane

The new control plane was built in Go as a plugin-driven system with 17 provider integrations (Infoblox, Kubernetes, Cisco UCS/Redfish, DCIM/CMDB, Grafana), a 51-field CMDB with an 11-state asset lifecycle, Dex OIDC RBAC, and multi-datacenter federation. Bare-metal provisioning was automated end to end with GitLab CI, Tinkerbell, Cisco UCS service profiles, and Pure Storage — HMAC-signed callbacks and idempotency guarantees brought p95 provisioning under 30 minutes. Delivery ran on GitOps (GitLab CI, Kaniko, Harbor, ArgoCD) with distroless builds, Vault-agent secret injection, CloudNativePG HA, and Prometheus/Grafana observability.

Engagement Results

The platform team adopted the 75-finding remediation roadmap; the Go control plane replaced ad-hoc automation with an auditable, RBAC-governed system; bare-metal provisioning reached sub-30-minute p95; and Kubernetes across multiple data centers was consolidated onto bare-metal Rancher RKE2 and Harvester. Ownership — runbooks, documentation, and GitOps workflows — was transferred to the customer's team. Detailed findings, architecture, and metrics are available to qualified buyers under NDA.

About THNKBIG

THNKBIG is an Austin, TX-based Platform Engineering consultancy specializing in Kubernetes, cloud-native, and production AI infrastructure. We build around how your business actually runs and transfer ownership to your team. Specialties: platform engineering, on-prem and cloud Kubernetes, software supply-chain security, infrastructure automation, and compliance readiness (FedRAMP, HIPAA, SOC 2, PCI, ATO).

Industry Context

Sector-Specific Challenges

Technology companies must maintain development velocity while ensuring platform reliability and security. These organizations face challenges in scaling microservices architectures, managing complex deployment pipelines, and maintaining developer productivity without compromising production stability.

Technical Considerations

Technology infrastructure requires support for continuous deployment practices, comprehensive observability across distributed systems, feature flag management, and chaos engineering capabilities. Systems must enable rapid experimentation while maintaining service level objectives.

Regulatory Environment

Technology companies typically must demonstrate SOC 2 Type II compliance for enterprise customers, GDPR compliance for European users, and often industry-specific certifications based on customer requirements.

Frequently Asked Questions

How do you approach client engagements?

Every engagement begins with a thorough discovery phase to understand your current state, business objectives, and constraints. We develop tailored recommendations rather than applying one-size-fits-all solutions. Our consultants work alongside your team to transfer knowledge and build sustainable capabilities. We measure success by business outcomes, not just technical deliverables.

What ROI can we expect from this type of engagement?

Organizations typically see significant improvements across multiple dimensions. Common outcomes include 50-80% reduction in deployment time, 30-50% decrease in infrastructure costs, 60-90% reduction in incident resolution time, and substantial improvements in developer productivity. The specific ROI depends on your starting point and investment level, which we help quantify during the assessment phase.

Related Solutions

This case study demonstrates our expertise in the following service areas. Learn more about how we can help your organization achieve similar results.

Cloud Complexity is a Problem — Until You Have the Right Team

From compliance automation to Kubernetes optimization, we help enterprises transform infrastructure into a competitive advantage.

Talk to a Cloud Expert