SYSTEM Cited by 9 sources
Amazon ECS (Elastic Container Service)¶
Amazon Elastic Container Service (Amazon ECS) is AWS's proprietary container-orchestration service, the substrate for both App Mesh (deprecated) and ECS Service Connect (current). Compared to Kubernetes (which AWS also offers as EKS), ECS is AWS-native with tighter IAM / VPC / ALB integration and a simpler abstraction model.
Stub page — minimal viable for the App Mesh discontinuation ingest. Expand on future ECS-internals sources.
Core abstractions¶
- Task Definition — template for a container group (image, resources, env, IAM role, networking mode).
- Task — running instance of a Task Definition; the unit of compute.
- Service — long-running managed group of identical Tasks with desired-count, health, and replacement policy.
- Cluster — logical group of Services / Tasks.
Tasks can run on EC2 capacity (customer-managed instances) or on Fargate (serverless, AWS-managed compute).
Role in service-mesh story¶
The ECS Service is the atomic unit for mesh membership:
- In App Mesh, each Service's Task Definition includes a self-managed Envoy sidecar container.
- In Service Connect, each Service's Task Definition cannot also be in App Mesh — the mesh membership is exclusive, which is why migration is forced to blue/green recreate.
Related¶
- systems/aws-app-mesh — deprecated sidecar-mesh layer above ECS
- systems/aws-ecs-service-connect — current managed mesh layer
- systems/smartprominence-ai-orchestrator — ECS Fargate as the stateless agent fleet (Terraform-provisioned task defs) in MHK's agentic orchestrator.
Seen in¶
-
sources/2026-10-01-aws-accelerating-airline-retailing-innovation-datalex-modernization — ECS/Fargate as the modernization target runtime. Datalex's Spring Boot / Java 21 microservices (extracted from a Java 8 / EJB2 n-tier system via a Strangler Fig business service proxy) run as Docker containers on ECS Fargate — multi-AZ, CPU/memory auto scaling — explicitly to avoid the operational overhead of managing EC2 instances running JBOSS. The DevSecOps pipeline uses native ECS rollback plus blue/green and canary deployment strategies for zero-downtime releases.
-
sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift — ECS as the compute tier evacuated in a multi-day AZ drill. Shifting at the ALB layer with ARC Zonal Shift stops new traffic to the evacuated AZ, but existing tasks there keep running — a complete evacuation also updates the ECS service
networkConfigurationto exclude the evacuated AZ's subnets (no new task placement there) and optionally setsdesired-countto an N-1 value. Load-bearing tuning: ECSstopTimeout: 55s(just below the 60s ALB deregistration delay) so tasks finish in-flight requests before force-stop, avoiding 502s. The post stresses pre-scaling for N-1 (static stability: ~+50% compute in a 3-AZ env) rather than reactive scaling, and that the service-update/count changes are control-plane operations to perform before a drill, not during a real impairment. (Source: sources/2026-09-30-aws-running-multi-day-az-evacuation-drills-with-arc-zonal-shift) - sources/2026-09-30-aws-how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock — ECS Fargate as the independently-scalable substrate for stateless agents in a controller-agent workflow, with task definitions provisioned by a config-driven dynamic agent registry. MHK's SmartProminence AI Orchestrator runs its orchestration core, workflow controllers, and LLM processing agents as ECS Fargate services; agents scale horizontally off SQS queue depth so agent processing scales independently of workflow logic. When a new agent type is registered through the core's API, Terraform auto-provisions the SQS queues, IAM roles, and ECS task definitions — the agent becomes available for workflow steps with no new compliance certification (self-service infrastructure + stateless compute). (Source: sources/2026-09-30-aws-how-mhk-built-a-hipaa-eligible-agentic-ai-solution-on-amazon-bedrock)
- sources/2026-09-10-aws-building-resilient-real-time-streaming-workers-with-amazon-dynamodb-leases
— ECS (on Fargate) as the worker-fleet substrate for a
stateful WebSocket coordination system. Two ECS behaviors are load-bearing:
(1) rolling deployments terminate old Tasks and start new ones — each terminated
Task drops its connections, which is exactly why the lease
pattern exists; (2) SIGTERM on Task stop (rolling deploy or scale-in) gives the
worker its graceful-shutdown window to release leases immediately (set
lease_expires_at_ms = 0) so peers reclaim connections without waiting for expiry. Scale-in is safe because of the lease pattern. Application Auto Scaling target-tracking drives Task count off a custom CloudWatchActiveConnectionsmetric (e.g. 700 conns/task target; 2 min scale-out / 15 min scale-in cooldowns). The post notes the lease logic is compute-agnostic — the same pattern works on EKS or EC2 Auto Scaling groups. (Source: sources/2026-09-10-aws-building-resilient-real-time-streaming-workers-with-amazon-dynamodb-leases) - sources/2026-05-12-aws-building-hybrid-multi-tenant-architecture-for-stateful-services
— canonical wiki instance of ECS cluster as the tenant-isolation
boundary for a stateful multi-tenant service. AWS's ad-serving
platform runs one dedicated ECS cluster per tenant inside shared
AWS accounts; each cluster loads only its tenant's in-memory
state, eliminating cross-tenant heap sharing. Canonical
dedicated-ECS-cluster-
per-tenant pattern with naming convention
(
tier-1-cell-1-ig-1-tenant-a), task-definitionTENANT_IDenv var propagation, and ECS-task-per-service 5,000 ceiling applying per tenant because cluster is single-tenant. Capacity math (up to 5 clusters per tenant, up to 100 ECS clusters per infra group) and the cluster- level tenant isolation framing first canonicalised on this ingest. - sources/2025-01-18-aws-app-mesh-discontinuation-service-connect-migration — the substrate under both meshes; the ECS Service's exclusive mesh-membership constraint is load-bearing for migration.
- sources/2024-08-08-figma-migrated-onto-k8s-in-less-than-12-months — ECS as the origin substrate in Figma's 12-month migration to EKS. Figma enumerates ECS limitations that drove the move: no StatefulSets (had to write custom etcd cluster-membership code on ECS, "fragile and hard to maintain"); no Helm support (OSS like systems/temporal required hand-porting into systems/terraform); poor graceful-node-drain on ECS-on-EC2 vs EKS cordon-and-drain; limited auto-scaling vs CNCF systems/keda + systems/karpenter; missing service-mesh off-the-shelf options; expected slower investment vs vendor-agnostic Kubernetes.
- sources/2026-04-08-aws-build-a-multi-tenant-configuration-system-with-tagged-storage-patterns — ECS tasks on Fargate in private subnets are the substrate for the NestJS gRPC Order Service + Config Service in the multi-tenant tagged-storage architecture; ECS Service as the registration unit discovered by Cloud Map for the event-driven refresh Lambda's gRPC fan-out.
- — ECS as the non-Kubernetes substrate that Zalando Payments' Load Test Conductor scales in lockstep with Kubernetes during an end-to-end load-test run. One declarative load-test API call scales applications across two Kubernetes node pools and an ECS cluster simultaneously — an instance of multi-substrate parallel orchestration from a single control plane for a heterogeneous microservice landscape.
Bosch L.OS tracking connector¶
Bosch L.OS runs its central Tracking Connector on ECS with Fargate. This shared orchestration layer standardizes provider protocols, routes requests, aggregates discovery acknowledgments, manages sessions, and handles errors; provider-specific Lambda adapters remain outside the connector. (Source: sources/2026-08-14-aws-serverless-vehicle-tracking-at-scale-bosch-los-on-aws)