Hedronite · Ops Lesson · 01-Earth-DevOps / Kubernetes · Wed 2026-09-30

EKS Fargate profiles — immutable selectors and a Rust coverage census

A Fargate profile is a create-time contract. Replace it to change it.

Lesson Class: Ops (DevOps + Kubernetes + EKS Fargate)
Topic: T2 Ops scripting
Cloud Referent: EKS Fargate profiles · selectors · alphanumeric pick
Automation: cargo script · aws-sdk-eks · kube 2 · Nix writeBashBin
Paired Dev: Async kube-rs LabelSelector Fargate census
Paired Cert: CKA HorizontalPodAutoscaler
Selector Contract
Namespace required. Labels AND. Five max. Private subnets.
Replace, do not patch
Create new profile, then delete old.
Alphanumeric First
Multiple matches: lowest profile name wins.
If two profiles match, the name that sorts first owns the pod.

<!-- hal:authoritative:yaml -->

A Fargate profile is a create-time contract. Change the contract by replacing the profile, not by editing it.

§I. Frame

On Amazon EKS, a Fargate profile tells the control plane which pods should run on Fargate instead of on EC2 nodes. The profile names a Pod execution role, one or more private subnets, and up to five selectors. Every selector must include a Kubernetes namespace. Labels on the selector are optional. If labels are present, a pod must carry every key-value pair. If labels are absent, every pod in that namespace is a candidate.

The EKS userguide states the hard edges in plain language. Profiles cannot be changed in place. You create a replacement profile, wait for it to become active, then delete the old one. While any profile in the cluster is DELETING, you cannot create another. When a profile is deleted, pods that were using it stop and move to Pending until another matching profile exists or you change the workload.

That is today's Ops surface. Not GKE surge upgrades (09-27), not AKS Disk CSI (09-24), not Pod Identity associations (09-21), not Gateway API (09-18). Access Entries stay for a later EKS fire.

The problem for today: describe the selector contract correctly, then build a read-only Rust census that lists every profile's selectors and flags pods whose Fargate label or selector match does not line up.

§II. Selector Contract

Selector Contract (named technique). Treat the profile as four facts you can verify without guessing:

  1. Namespace is required. A selector with no namespace is invalid. Wildcards * and ? are allowed in namespace and label keys/values (prod*, key?).
  2. Labels are AND. Every listed label must be present on the pod. Omitting labels means "whole namespace".
  3. Five selectors max per profile. Need a sixth namespace? Add another profile.
  4. Private subnets only. Fargate pods do not get public IPs. Subnets need NAT to reach ECR and the API, not a direct Internet Gateway route.

When two profiles both match a pod, EKS picks the profile whose name sorts first alphanumerically, then stamps eks.amazonaws.com/fargate-profile=<that-name> on the pod. You can force a named profile by setting that label yourself, but the pod must still match a selector in that profile.

eksctl create fargateprofile \
  --cluster prod-a \
  --name batch-a \
  --namespace batch \
  --labels app=etl

Patching --labels later is not an API. Create batch-b with the new selectors, confirm Active, delete batch-a.

§III. Coverage gaps worth hunting

Three failure shapes show up in production reviews:

SymptomUsual cause
Pod Pending after profile deleteOld profile gone; no replacement selector yet
Pod on EC2 despite "we use Fargate"Namespace or label mismatch; selector never matched
Wrong profile name on the podTwo profiles matched; alphanumeric pick won

A census that only prints profile names misses the third row. You need Describe on each profile (selectors) and a Pod list (labels + namespace).

§IV. Rust census (cargo script + kube-rs)

eks-fargate-coverage-census.rs is a cargo script. It takes one argument, the cluster name.

  1. aws-sdk-eks lists Fargate profile names, then describes each for selectors.
  2. kube-rs lists every pod (read-only).
  3. For each pod it classifies: ON_FARGATE (label + known profile), ORPHAN_LABEL (label without a described profile), SELECTOR_HIT_NO_LABEL (would match on reschedule but no Fargate label yet), or ignore.

Core match (same rules as section II):

fn pod_matches(pod: &Pod, sel: &Selector) -> bool {
    let ns = pod.metadata.namespace.as_deref().unwrap_or("");
    if ns != sel.namespace {
        return false;
    }
    let empty = BTreeMap::new();
    let labels = pod.metadata.labels.as_ref().unwrap_or(&empty);
    sel.labels.iter().all(|(k, v)| labels.get(k) == Some(v))
}

Alphanumeric First (named technique). When the pod lacks an explicit eks.amazonaws.com/fargate-profile label, collect every profile with at least one matching selector, sort by profile name, and take the first. That mirrors the documented EKS conflict rule so the census and the scheduler argue from the same order.

Illustrative output (not a live cluster run this fire):

PROFILE batch-a selectors=batch/[app=etl]
PROFILE kube-dns selectors=kube-system/*
ON_FARGATE kube-system/coredns-… profile=kube-dns
SELECTOR_HIT_NO_LABEL batch/etl-worker-0 would_use=batch-a phase=Some("Running")
profiles=2 on_fargate=2 gaps=1

SELECTOR_HIT_NO_LABEL on a Running pod usually means the workload landed on EC2 before the profile existed, or the labels were added after schedule. Recycle the pod (delete and let the controller recreate) after the profile is Active.

What was checked, and what was not. The script targets aws-sdk-eks 1.x, kube 2, k8s-openapi 0.26 with v1_34, and tokio. The lab Mac runs cargo check on a scratch crate from the frontmatter after ship. The live -Zscript path needs AWS credentials for Describe and a reachable kubeconfig for the Pod list. Neither is exercised against a real cluster this fire (no local container runtime, no k8s mutation). Sample lines above are illustrative.

§V. Wrap it with Nix

eks-fargate-coverage-census.nix:

{ pkgs ? import <nixpkgs> { } }:
pkgs.writers.writeBashBin "eks-fargate-coverage-census" ''
  exec cargo +nightly -Zscript ${./eks-fargate-coverage-census.rs} "$@"
''

Same pattern as the 09-27 GKE PDB census wrapper: a stable command name, still dependent on the caller's cargo toolchain.

§VI. What not to do

  1. Editing selectors on an existing profile in Terraform or the console and expecting an in-place update. Replacement is the contract.
  2. Creating a second profile while another is DELETING.
  3. Putting Fargate pods in public subnets.
  4. Assuming affinity or topologySpreadConstraints will place Fargate pods. EKS Fargate ignores those rules; spread comes from subnet list design (often one subnet per profile when even spread matters).
  5. Giving the census write verbs on pods or any Fargate mutate IAM.

§VII. Close instruction

Write one profile with two selectors: namespace batch with label app=etl, and namespace batch with label app=report. Say which profile name wins if a second profile batch-0 also matches app=etl, and what label the pod should show. Then name the exact replace sequence when you need to add a third label key.

Related