Rust over k8s-openapi — a drain preflight
kubectl drain makes one decision per pod. Make the same decisions offline, in types.
<!-- hal:authoritative:yaml -->
*kubectl drain makes one decision per pod. Make the same decisions offline, in types, before the drain makes them for real.*
§I. Frame
kubectl drain sorts every pod on a node into a small number of outcomes, and its flags name them. From kubectl drain --help (kubectl 1.37.0 on the lab Mac): --ignore-daemonsets to skip DaemonSet pods, --force to "continue even if there are pods that do not declare a controller", and --delete-emptydir-data to "continue even if there are pods using emptyDir". Past those gates, each remaining pod gets an Eviction request, and the API server answers 200, or 429 when a PodDisruptionBudget has no headroom, or 500 when more than one budget covers the pod.
The Ops lesson in this trio asks the live cluster which budgets will stall a GKE pool. This lesson answers a narrower question with no cluster at all: given kubectl get pods -A -o json and kubectl get pdb -A -o json saved to disk, what will kubectl drain gke-prod-pool-a-1 do to each pod?
The crate drain-preflight sits in this bundle under pkg/. It depends on k8s-openapi for the Kubernetes types and on serde_json to read them, with no async and no client.
§II. The kubectl List envelope
The first version read the file as k8s_openapi::List<Pod>. Every test failed with the same message:
Error("invalid value: string \"List\", expected PodList", line: 3, column: 16)
kubectl get ... -o json does not return the API server's typed list. It returns a generic envelope with "kind": "List", and k8s-openapi's List<T> checks the kind field against PodList and refuses. That check is correct for a response from the API server. It is wrong for a file a human saved from kubectl.
Envelope Strip (named technique). Deserialize the envelope you actually have, keep only what you need, and let the typed items do the checking:
#[derive(Deserialize)]
struct KubectlList<T> {
items: Vec<T>,
}
pub fn items<T: DeserializeOwned>(json: &str) -> Result<Vec<T>, serde_json::Error> {
let list: KubectlList<T> = serde_json::from_str(json)?;
Ok(list.items)
}
The generic parameter and its DeserializeOwned bound are ch10 material. One function reads both files: items::<Pod> and items::<PodDisruptionBudget>. The envelope is loose on purpose, and the items stay strict, so a malformed pod still fails to parse. A test now pins the original failure, so nobody "fixes" the code back to List<Pod>:
#[test]
fn typed_list_rejects_kubectl_output() {
let typed: Result<k8s_openapi::List<Pod>, _> =
serde_json::from_str(include_str!("../fixture/pods.json"));
assert!(typed.is_err());
}
§III. One enum, one verdict per pod
#[derive(Debug, PartialEq)]
pub enum Verdict {
DaemonSet, // skipped with --ignore-daemonsets
Mirror, // static pod; drain never touches it
Unmanaged, // no controller: drain refuses unless --force
LocalData, // emptyDir: refuses unless --delete-emptydir-data
Blocked(String), // one PDB, zero headroom: eviction 429
MultiPdb(Vec<String>),// two or more PDBs: eviction 500
Evictable,
}
The code in pkg/src/lib.rs carries these notes as doc comments. classify checks the gates in a fixed order and returns at the first one that fires. The owner check shows why k8s-openapi types are worth the ceremony:
let controller = meta
.owner_references
.iter()
.flatten()
.find(|o| o.controller == Some(true));
match controller {
Some(o) if o.kind == "DaemonSet" => return Verdict::DaemonSet,
None => return Verdict::Unmanaged,
Some(_) => {}
}
owner_references is Option<Vec<OwnerReference>>, and controller inside it is Option<bool>. .iter().flatten() turns "maybe a list" into "zero or more items" without an unwrap. Comparing against Some(true) treats a missing flag as "not the controller", which matches the API's meaning. The match must cover None, and the compiler enforces it: TRPL ch6 shows the same non-exhaustive patterns: None not covered error. A pod with owners but no controller owner lands in Unmanaged, which is what --force means by "pods that do not declare a controller".
§IV. Matching selectors by hand
The Ops census hands the PDB selector to the API server. Offline there is no server, so labels_match implements the rules itself:
fn labels_match(sel: &LabelSelector, labels: &BTreeMap<String, String>) -> bool {
let by_label = sel
.match_labels
.iter()
.flatten()
.all(|(k, v)| labels.get(k) == Some(v));
let by_expr = sel.match_expressions.iter().flatten().all(|req| {
let have = labels.get(&req.key);
let values = req.values.as_deref().unwrap_or(&[]);
match req.operator.as_str() {
"In" => have.is_some_and(|v| values.contains(v)),
"NotIn" => have.is_none_or(|v| !values.contains(v)),
"Exists" => have.is_some(),
"DoesNotExist" => have.is_none(),
_ => false, // an operator we do not know selects nothing
}
});
by_label && by_expr
}
Two decisions are visible. NotIn matches a pod that lacks the key entirely, which is the Kubernetes rule and the easiest one to get wrong. An unknown operator selects nothing, so a typo can only make the preflight report fewer covered pods, never invent coverage. covering then filters PDBs to the pod's own namespace before matching, and returns Vec<&'a PodDisruptionBudget>: references into the caller's slice, with the lifetime saying so (ch10).
§V. The run
cargo test passed 4 on the lab Mac (cargo 1.96.0, k8s-openapi 0.26.1), and cargo clippy was clean. The fixture puts nine pods and five PDBs in the kubectl envelope. Two pods are filtered before classification: a Succeeded Job pod on the node and a web replica on another node. One PDB, staging/api-pdb, has the same selector as prod/api-pdb and must not match the prod pod. The run:
$ cargo run -q -- gke-prod-pool-a-1 fixture/pods.json fixture/pdbs.json
skip kube-system/fluentbit-x2k9p needs --ignore-daemonsets
skip kube-system/kube-proxy-gke-prod-pool-a-1 static pod, drain ignores
STOP prod/debug-shell no controller; --force deletes it for good
STOP prod/cache-0 emptyDir; --delete-emptydir-data loses it
WAIT prod/api-0 api-pdb allows 0 disruptions; eviction 429
ok prod/web-7d9f-abcde evicts now
STOP prod/worker-5c6b-zzzzz 2 PDBs match (worker-pdb,tier-backend-pdb); eviction 500
$ echo $?
3
The worker row is the one a human review misses. worker-pdb selects app=worker, and tier-backend-pdb selects tier In (backend). Each budget has headroom on its own, and the eviction still fails with a 500, because the API refuses to guess which budget governs. Exit codes follow the ch12 pattern: 0 when clean, 3 when any pod would stop or wait, 64 for bad usage, 2 for unreadable input.
§VI. What not to do
- Parsing kubectl output as
List<Pod>and blaming the fixture when it fails. - Calling
unwraponowner_referencesorstatus. Both are legitimately absent on real objects. - Treating
NotInas "has the key with another value". - Matching PDBs across namespaces. A budget only covers pods in its own namespace.
- Pulling in kube-rs here. The live client is async, and the Ops cargo script already owns that job.
§VII. Close instruction
Add a test pod on the node owned by a ReplicaSet with tier=backend and app=queue, and predict its verdict before running cargo test. Then add a Verdict::Finished variant for Succeeded and Failed pods instead of filtering them out, and let the compiler show you every match that has to change.
Related
- Ops: GKE surge upgrades and PDB stalls (same trio)
- Cert: CKS upgrade, version skew, and drain (same trio)
- Prior Dev: serde over terraform plan JSON