Skip to content

CONCEPT Cited by 3 sources

Self-service infrastructure

The discipline of enabling engineers to provision, configure, and tear down infrastructure environments without manual ops intervention — while maintaining central governance, cost attribution, and security controls.

Definition

Self-service infrastructure shifts the provisioning burden from platform/ops teams to the engineers who need the resources, mediated by a governed platform layer. The key tension: engineers need speed and autonomy; organisations need governance, cost control, and security.

Structural Properties

  1. Template-governed — provisioning paths are pre-vetted and hardened; engineers choose from a catalog, not raw IaC
  2. Identity-aware — every resource is attributed to an owner with purpose metadata
  3. Lifecycle-managed — resources have TTLs, expiration notifications, and automatic cleanup
  4. Observable — every provisioning action is transparent, auditable, and attributable
  5. Multi-tenant safe — isolation between provisioned environments prevents interference

Failure Modes of Shared Environments (That Self-Service Solves)

  • Coordination overhead at scale — multiple engineers in one environment interfere during critical demos
  • Platform-limit contention — shared catalogs/instances hit hard limits
  • Cost attribution opacity — ownership unclear as usage increases
  • Observability gaps — tracing what happened, when, and why requires manual investigation

Known Instances

System Scale Approach
systems/databricks-fevm 5,000+ users, 2,600+ deployments Use-case-based provisioning via Databricks Apps + Terraform + MCP
Stripe Projects Agent-provisioned accounts CLI-driven account+service provisioning for AI agents
Equinix North Star (EKS) Multiple business orgs, 100% adoption Cloud-ops owns a governed shared platform; app teams self-serve pre-configured namespaces + shared data/CI-CD/observability — no cluster provisioning

Seen In

  • sources/2026-07-23-databricks-self-serve-infrastructure-vending-machine — FEVM handles the full lifecycle (provision → use → expire → cleanup) for 5,000+ field engineers across 3 clouds
  • sources/2026-09-17-aws-how-equinix-cut-operational-overhead-with-a-shared-services — Equinix escapes "operational sprawl" from decentralized per-team cluster ownership by moving to a centrally-governed shared-services platform on EKS. The cloud operations team owns/governs the platform; application teams onboard into pre-configured namespaces with access to shared data services, centralized GitHub-Runner CI/CD, and observability — deploying independently "without provisioning or managing cluster infrastructure." Reported: weeks → days onboarding and 4x deployment frequency, the canonical self-service payoff (governance + speed).
  • sources/2026-09-30-aws-how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock — Config-driven agent onboarding with inherited compliance. MHK's SmartProminence AI Orchestrator has a dynamic agent registry: to add an agent type a developer defines prompt templates, input/output schemas, and model selection, then registers it through the orchestration core's API — and Terraform auto-provisions the supporting SQS queues, IAM roles, and ECS task definitions. The agent is immediately available for workflow steps with no new compliance certification, because it runs inside the already-certified orchestrator (per-client encryption, audit logging, token-based access, data minimization are inherited). The self-service payoff applied to agentic AI features: engineers focus on prompt + workflow logic, not infra scaffolding or per-feature HIPAA/SOC 2 attestation — the mechanism behind the reported 3–6-months → ~2-weeks deployment collapse.

Merged aliases

  • infrastructure-vending-machine
Last updated · 766 distilled / 2,225 read