CONCEPT Cited by 3 sources
Self-service infrastructure¶
The discipline of enabling engineers to provision, configure, and tear down infrastructure environments without manual ops intervention — while maintaining central governance, cost attribution, and security controls.
Definition¶
Self-service infrastructure shifts the provisioning burden from platform/ops teams to the engineers who need the resources, mediated by a governed platform layer. The key tension: engineers need speed and autonomy; organisations need governance, cost control, and security.
Structural Properties¶
- Template-governed — provisioning paths are pre-vetted and hardened; engineers choose from a catalog, not raw IaC
- Identity-aware — every resource is attributed to an owner with purpose metadata
- Lifecycle-managed — resources have TTLs, expiration notifications, and automatic cleanup
- Observable — every provisioning action is transparent, auditable, and attributable
- Multi-tenant safe — isolation between provisioned environments prevents interference
Failure Modes of Shared Environments (That Self-Service Solves)¶
- Coordination overhead at scale — multiple engineers in one environment interfere during critical demos
- Platform-limit contention — shared catalogs/instances hit hard limits
- Cost attribution opacity — ownership unclear as usage increases
- Observability gaps — tracing what happened, when, and why requires manual investigation
Known Instances¶
| System | Scale | Approach |
|---|---|---|
| systems/databricks-fevm | 5,000+ users, 2,600+ deployments | Use-case-based provisioning via Databricks Apps + Terraform + MCP |
| Stripe Projects | Agent-provisioned accounts | CLI-driven account+service provisioning for AI agents |
| Equinix North Star (EKS) | Multiple business orgs, 100% adoption | Cloud-ops owns a governed shared platform; app teams self-serve pre-configured namespaces + shared data/CI-CD/observability — no cluster provisioning |
Seen In¶
- sources/2026-07-23-databricks-self-serve-infrastructure-vending-machine — FEVM handles the full lifecycle (provision → use → expire → cleanup) for 5,000+ field engineers across 3 clouds
- sources/2026-09-17-aws-how-equinix-cut-operational-overhead-with-a-shared-services — Equinix escapes "operational sprawl" from decentralized per-team cluster ownership by moving to a centrally-governed shared-services platform on EKS. The cloud operations team owns/governs the platform; application teams onboard into pre-configured namespaces with access to shared data services, centralized GitHub-Runner CI/CD, and observability — deploying independently "without provisioning or managing cluster infrastructure." Reported: weeks → days onboarding and 4x deployment frequency, the canonical self-service payoff (governance + speed).
- sources/2026-09-30-aws-how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock — Config-driven agent onboarding with inherited compliance. MHK's SmartProminence AI Orchestrator has a dynamic agent registry: to add an agent type a developer defines prompt templates, input/output schemas, and model selection, then registers it through the orchestration core's API — and Terraform auto-provisions the supporting SQS queues, IAM roles, and ECS task definitions. The agent is immediately available for workflow steps with no new compliance certification, because it runs inside the already-certified orchestrator (per-client encryption, audit logging, token-based access, data minimization are inherited). The self-service payoff applied to agentic AI features: engineers focus on prompt + workflow logic, not infra scaffolding or per-feature HIPAA/SOC 2 attestation — the mechanism behind the reported 3–6-months → ~2-weeks deployment collapse.
Merged aliases¶
infrastructure-vending-machine