Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
43 commits
Select commit Hold shift + click to select a range
8292dac
test(operator): define CRD lifecycle contract
Jul 21, 2026
73f1d74
feat(operator): enforce CRD lifecycle gates
Jul 21, 2026
0cbf2b2
fix(operator): close lifecycle contract test
Jul 21, 2026
c6b333c
fix(operator): use structural fixture round trips
Jul 21, 2026
2747843
fix(operator): assert default without new modules
Jul 21, 2026
3e84b06
test(operator): expose CRD lifecycle blockers
Jul 21, 2026
53d7c1c
fix(operator): validate proposed CRD semantics
Jul 21, 2026
20e2384
test(operator): account for local preflight
Jul 21, 2026
721c84c
ci(operator): emit validator module metadata
Jul 21, 2026
7b1affa
ci(operator): freeze validator dependencies
Jul 21, 2026
1b84fa7
ci(operator): preserve ordinary Go gate
Jul 21, 2026
e27f3f9
test(operator): expose transition CEL gap
Jul 21, 2026
8bd18a1
fix(operator): validate unchanged CEL transitions
Jul 21, 2026
1a207e1
test(operator): expose live CR narrowing gap
Jul 22, 2026
de9bf6f
fix(operator): validate every live CR before apply
Jul 22, 2026
2ebbf13
test(operator): require dependency event convergence
Jul 21, 2026
f91517b
feat(operator): reconcile dependency changes
Jul 21, 2026
72c6a19
test(operator): require host policy convergence
Jul 21, 2026
bb851a8
fix(operator): fail closed on host storage drift
Jul 21, 2026
4ccae52
fix(operator): require explicit workspace storage class
Jul 21, 2026
30752e2
test(operator): reject stale dependency authority
Jul 21, 2026
95da911
fix(operator): revoke untrusted dependency authority
Jul 21, 2026
d954667
test(operator): revoke terminal storage authority
Jul 21, 2026
187baf2
fix(operator): clear terminal storage authority
Jul 21, 2026
75762da
test(operator): stop owned siblings on conflict
Jul 21, 2026
55539e5
fix(operator): stop owned siblings on conflict
Jul 21, 2026
764f460
test(operator): preserve dependency status on conflicts
Jul 21, 2026
e24b6dc
fix(operator): preserve ready dependencies on conflicts
Jul 21, 2026
96ccd96
test(operator): cover cleanup and create races
Jul 21, 2026
2a8b7ae
fix(operator): close child ownership races
Jul 21, 2026
e366c52
test(operator): model stale cache after create race
Jul 21, 2026
4916c10
test(operator): cover workspace PVC create race
Jul 21, 2026
db95dd6
fix(operator): validate authoritative create winners
Jul 21, 2026
ee5d009
test(operator): cover authoritative cleanup policy
Jul 21, 2026
d9ef139
fix(operator): precondition authoritative cleanup
Jul 21, 2026
32d4a04
test(operator): cover split-reader dependency cleanup
Jul 21, 2026
a10e70b
fix(operator): authoritatively clean revoked sessions
Jul 21, 2026
3d09b4a
test(operator): expose stale workspace session cache
Jul 21, 2026
d428c0a
fix(operator): list workspace sessions authoritatively
Jul 21, 2026
4a7cdf6
test(operator): reject stale PVC authority
Jul 22, 2026
b98470d
fix(operator): revalidate PVC authority
Jul 22, 2026
ece837c
test(operator): gate existing pods on PVC authority
Jul 22, 2026
20e9878
fix(operator): gate child repairs on PVC authority
Jul 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -294,6 +294,8 @@ jobs:

- name: Run operator tests
working-directory: packages/cluster-operator
env:
GOFLAGS: -mod=readonly
run: go test ./...

- name: Lint cluster chart
Expand Down
4 changes: 2 additions & 2 deletions .woodpecker.yml
Original file line number Diff line number Diff line change
Expand Up @@ -250,9 +250,9 @@ steps:
image: mirror.gcr.io/library/golang:1.25.1-bookworm@sha256:c423747fbd96fd8f0b1102d947f51f9b266060217478e5f9bf86f145969562ee
directory: packages/cluster-operator
commands:
- GOMAXPROCS=1 GOFLAGS=-p=1 go test ./api/... ./controllers/... ./cmd/...
- GOMAXPROCS=1 GOFLAGS='-mod=readonly -p=1' go test ./api/... ./controllers/... ./cmd/...
- mkdir -p ../../artifacts/cluster-proof
- CGO_ENABLED=0 GOMAXPROCS=1 GOFLAGS=-p=1 go test -c ./charttests -o ../../artifacts/cluster-proof/chart-contract.test
- CGO_ENABLED=0 GOMAXPROCS=1 GOFLAGS='-mod=readonly -p=1' go test -c ./charttests -o ../../artifacts/cluster-proof/chart-contract.test
backend_options:
kubernetes:
serviceAccountName: woodpecker-ci-untrusted
Expand Down
98 changes: 82 additions & 16 deletions docs/CLUSTER_OPERATOR.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,13 +23,15 @@ That declaration does not make a non-RWX backend safe. Missing classes, classes

## Install

Installing with defaults creates no controller, gateway, session workload, RBAC, Secret, or network policy:
Installing with defaults creates no controller, gateway, session workload, RBAC, Secret, or network policy. Use the lifecycle runner even for a fresh install so existing live objects and compatibility fixtures are checked against the proposed schemas before the CRDs are server-validated, established, and storage-version-checked and before Helm runs:

```sh
helm install t4-cluster deploy/charts/t4-cluster --namespace t4-system --create-namespace
scripts/cluster-ci/crd-lifecycle.sh install -- \
helm install t4-cluster deploy/charts/t4-cluster \
--namespace t4-system --create-namespace --skip-crds
```

Helm processes files in `crds/` separately. Use `--skip-crds` if CRD lifecycle is administered independently. To enable the control plane, provide a private values file or a deployment controller values source:
Helm processes files in `crds/` separately on a direct install, but does not upgrade or delete them later. The release procedure therefore administers CRDs independently and always passes `--skip-crds` to Helm. To enable the control plane, provide a private values file or a deployment controller values source:

```yaml
enabled: true
Expand Down Expand Up @@ -112,11 +114,7 @@ The controller uses uncached, exact-name `get` permission for only that administ

Reusable credential projection is unsupported because OMP and arbitrary session tools share one workload security boundary. `allowUnauthenticated: true` with empty `credentialSecret` and `credentialKey` is therefore mandatory. The two credential fields remain only as cutover sentinels: Helm, the controller, and the image entrypoint all reject an old credential-mode configuration instead of exposing it. Before OMP starts, the image requires the pinned `auth_credentials` table to be empty, rejects the pinned secret settings (`auth.broker.token`, `hindsight.apiToken`, `searxng.token`, and `dev.autoqaPush.token`), rejects broker token and encrypted snapshot files, and fails closed on unknown schemas or linked profile paths. It never deletes or rewrites durable user state. Because Helm rejects legacy credential values before rendering the new Deployment, a failed upgrade leaves the prior release unchanged. Stop development sessions and clear any legacy values and credential state explicitly before adopting this pre-production boundary. Provider authentication must live behind a private model gateway that authorizes the session through infrastructure identity or network policy while presenting an `auth: none` endpoint inside the session's allowed route. Do not expose that endpoint to the Internet or a shared untrusted network. The chart remains provider- and model-neutral.

Install or enable with immutable digests:

```sh
helm upgrade --install t4-cluster deploy/charts/t4-cluster --namespace t4-system --values operator-values.yaml
```
Install a new release with immutable digests by adding `--values operator-values.yaml` to the lifecycle-runner `install` command above. For an existing release, use the lifecycle-runner `upgrade` procedure below. Do not use `helm upgrade --install`: install and upgrade have different compatibility preflights.

The controller always has two replicas and uses a Kubernetes Lease named from `t4-cluster-operator.cluster.t4.dev`; one replica reconciles at a time. The server defaults to three stateless replicas and supports a minimum of two. Its Deployment uses `maxUnavailable: 0`, a `minAvailable: 2` PDB, topology spread, anti-affinity, readiness draining, and an explicit `k3s-worker-02` exclusion. Session pods also exclude that node by default. Additional cluster-specific exclusions belong in deployment values, not this portable chart.

Expand Down Expand Up @@ -150,28 +148,96 @@ The session runtime verifies and builds the exact OMP tag `t4code-17.0.5-appserv

Session pods do not receive an automatically mounted ServiceAccount token. The explicit projected reviewer token can only create TokenReviews and is not resource-authorized. The namespace-local ConfigMap projection contains credential-free `auth: none` OMP models and credential-free settings; session Pods receive no provider Secret reference or reusable provider credential. All containers drop capabilities, disallow privilege escalation, use RuntimeDefault seccomp, and use read-only root filesystems. No per-session NodePort, LoadBalancer, host network, host PID, host display, or hostPath is created.

## Upgrade and rollback
## Upgrade, storage migration, and rollback

CRD changes must remain structural and additive. Helm installs files under `crds/` on first install but does not upgrade or delete them. Before every `helm upgrade`, an administrator must review and apply each newer CRD independently, wait for API discovery to serve the updated schemas, and only then upgrade workloads with the new immutable image digests:
The three namespaced CRDs have exactly one version: `cluster.t4.dev/v1alpha1`, with `served: true`, `storage: true`, a structural schema, and a status subresource. Changes in this version must remain additive: existing fields keep their meaning and validation, while new fields are optional or have safe defaults. The compatibility fixtures in `packages/cluster-operator/api/v1alpha1/testdata/compat/` are persisted old-object shapes, including status and finalizers, and must continue to validate and round-trip without losing any declared field.

Helm does not upgrade CRDs. For every workload upgrade, run the fail-closed lifecycle command:

```sh
kubectl apply --server-side -f deploy/charts/t4-cluster/crds/
kubectl wait --for=condition=Established crd/t4clusterhosts.cluster.t4.dev crd/t4workspaces.cluster.t4.dev crd/t4sessions.cluster.t4.dev
scripts/cluster-ci/crd-lifecycle.sh upgrade -- \
helm upgrade t4-cluster deploy/charts/t4-cluster \
--namespace t4-system --values operator-values.yaml --skip-crds
```

Do not rely on `helm upgrade` to change CRDs. If a CRD pre-upgrade apply fails validation, stop the upgrade and keep the currently deployed controller/server images; do not force-replace the CRD or remove stored versions.
The command performs this exact order:

1. Before any cluster access, validate every complete compatibility fixture (including `status`) against the proposed structural OpenAPI schema and CEL programs with the Kubernetes apiextensions validators. The check also rejects declared fields that the structural pruner would remove.
2. With `get` access to the three exact CRD definitions and `list` access to the corresponding namespaced T4 resources, detect which definitions already exist, enumerate every live CR for each existing definition across namespaces, and validate each complete object, including `status`, against the same proposed OpenAPI, CEL create, CEL unchanged-update, and pruning checks. A denied CRD read, denied or malformed object list, or incompatible object fails closed; an absent definition is safely empty on a fresh install.
3. Server-side dry-run the reviewed CRDs with strict validation. Upgrades stop before any mutation if a local proposed-schema validation, live-object enumeration, or CRD dry-run fails.
4. Apply the CRDs server-side without force conflicts.
5. Wait for all three CRDs to report `Established`.
6. Fetch the served `cluster.t4.dev/v1alpha1` OpenAPI v3 document three times and require every published resource schema to match the semantics generated from the proposed CRDs. Retained `Established=True` is not readiness when discovery still serves an old schema.
7. Server-side dry-run the compatibility fixtures against the converged admission path.
8. Require each CRD's `status.storedVersions` to be exactly `v1alpha1`.
9. Execute the supplied Helm command, which must use `--skip-crds`.

The corresponding administrative checks are executable independently:

```sh
helm upgrade t4-cluster deploy/charts/t4-cluster --namespace t4-system --values operator-values.yaml
(cd packages/cluster-operator && \
go run ./cmd/crd-preflight fixtures ../../deploy/charts/t4-cluster/crds api/v1alpha1/testdata/compat)
live_objects=$(mktemp)
trap 'rm -f "$live_objects"' EXIT
trap 'exit 129' HUP
trap 'exit 130' INT
trap 'exit 143' TERM
for resource in t4clusterhosts.cluster.t4.dev t4workspaces.cluster.t4.dev t4sessions.cluster.t4.dev; do
installed_crd=$(kubectl get "crd/$resource" --ignore-not-found -o name)
if [ -z "$installed_crd" ]; then
continue
fi
kubectl get "$resource" --all-namespaces -o json >"$live_objects"
(cd packages/cluster-operator && \
go run ./cmd/crd-preflight objects ../../deploy/charts/t4-cluster/crds <"$live_objects")
done
rm -f "$live_objects"
trap - EXIT HUP INT TERM
kubectl apply --server-side --dry-run=server --validate=strict \
--field-manager=t4-crd-lifecycle -f deploy/charts/t4-cluster/crds/
kubectl apply --server-side --validate=strict \
--field-manager=t4-crd-lifecycle -f deploy/charts/t4-cluster/crds/
kubectl wait --for=condition=Established --timeout=120s \
crd/t4clusterhosts.cluster.t4.dev \
crd/t4workspaces.cluster.t4.dev \
crd/t4sessions.cluster.t4.dev
for observation in 1 2 3; do
kubectl get --raw /openapi/v3/apis/cluster.t4.dev/v1alpha1 | \
(cd packages/cluster-operator && go run ./cmd/crd-preflight served ../../deploy/charts/t4-cluster/crds)
done
kubectl apply --dry-run=server --validate=strict --namespace default \
-f packages/cluster-operator/api/v1alpha1/testdata/compat/
for crd in t4clusterhosts.cluster.t4.dev t4workspaces.cluster.t4.dev t4sessions.cluster.t4.dev; do
test "$(kubectl get "crd/$crd" -o 'jsonpath={.status.storedVersions[*]}')" = v1alpha1
done
```

For an application rollback, retain the additive CRDs and use the previous known-compatible values and image digest set:
Do not rely on `helm upgrade` to change CRDs. Never use `kubectl replace --force`, `kubectl apply --force-conflicts`, delete/recreate a CRD, or alter `status.storedVersions` outside the verified migration sequence below. A preflight failure leaves the live CRDs, custom resources, controller/server workloads, and session workloads untouched. A failure after additive CRD apply but before Helm leaves the prior workloads running against the still-backward-compatible schema; investigate and rerun the gates rather than attempting CRD rollback.

### Future `v1beta1` conversion and storage procedure

There is no `v1beta1` API today. A future proposal is a separate incompatible lifecycle change and may proceed only through these gates:

1. Back up all three CRDs and every custom resource, record per-kind object counts, and prove a restore in an isolated cluster.
2. Review separate `v1beta1` schemas, defaults, conversion semantics, downgrade semantics, and old/new/old fixture round-trips. Every field represented in either version must survive conversion; conversion may not manufacture product or runtime state in status.
3. Deploy a highly available conversion webhook with a PodDisruptionBudget, strict TLS/service identity, readiness checks, failure policy, metrics, and alerts. Prove conversion availability before adding a served version to any CRD.
4. Add `v1beta1` as served but not storage, retain `v1alpha1` as served and storage, wait for `Established`, and prove reads and writes through both versions plus all compatibility fixtures. Do not advance while any controller, gateway, backup, or recovery tool cannot read both versions.
5. In a later reviewed CRD apply, set `v1beta1` storage to true and `v1alpha1` storage to false while keeping both served. Reconfirm conversion health, then rewrite every object through the Kubernetes API using an approved storage-version migrator; merely changing `spec.versions[*].storage` does not migrate stored objects.
6. Compare pre/post object counts and read every rewritten object through both served versions. Verify identity, declared spec and status fields, finalizers, and conversion round-trips. Stop on any missing object or lossy read.
7. Only after that verification, explicitly retire the old storage record through the CRD status subresource for each CRD, for example `kubectl patch customresourcedefinition t4sessions.cluster.t4.dev --subresource=status --type=json -p='[{"op":"replace","path":"/status/storedVersions","value":["v1beta1"]}]'`. Repeat for hosts and workspaces. This status update records completed storage migration; it does not perform migration.
8. Read each CRD back and require `status.storedVersions` to be exactly `[v1beta1]`. Keep `v1alpha1` served, the webhook available, the verified backup, and dual-read binaries for the entire rollback window. Removing the old served version is a later explicit change after the window expires.

If conversion or migration fails, stop forward rollout. Do not force-replace or downgrade CRDs. While both versions remain served and conversion is healthy, roll workloads back to the prior dual-read image set. If in-place compatibility cannot be proven, stop mutations and restore the verified backup under the reviewed recovery procedure before resuming service.

### Workload rollback

CRDs remain additive and installed. Roll back the known-compatible controller, server, T4 session runtime, and OMP digest set together:

```sh
helm rollback t4-cluster REVISION --namespace t4-system
```

Do not roll OMP independently of the T4 session runtime. Roll back the known-compatible T4/OMP image set together. The pinned authority boundary is not negotiated down.
Do not roll OMP independently of the T4 session runtime. The pinned authority boundary is not negotiated down.

## Uninstall

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
apiVersion: cluster.t4.dev/v1alpha1
kind: T4ClusterHost
metadata:
name: legacy-host
spec:
storageClassName: legacy-rwx
runtimeProfiles:
- default
- gui
ciProvider:
secretRef:
name: legacy-ci-token
configMapRef:
name: legacy-ci-config
allowedOrigins:
- https://t4.example.test
status:
observedGeneration: 3
conditions:
- type: Available
status: "True"
observedGeneration: 3
lastTransitionTime: "2026-01-15T12:00:00Z"
reason: Reconciled
message: Legacy host remains available
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
apiVersion: cluster.t4.dev/v1alpha1
kind: T4Session
metadata:
name: legacy-session
finalizers:
- cluster.t4.dev/session-cleanup
spec:
hostRef: legacy-host
workspaceRef: legacy-workspace
title: Legacy session
runtimeProfile: gui
initialPromptSecretRef:
name: legacy-initial-prompt
ci:
repositoryId: legacy/project
ref: refs/heads/main
commit: 0123456789abcdef0123456789abcdef01234567
status:
observedGeneration: 7
podName: legacy-session-pod
serviceName: legacy-session
phase: Running
conditions:
- type: Available
status: "True"
observedGeneration: 7
lastTransitionTime: "2026-01-15T12:00:00Z"
reason: RuntimeReady
message: Legacy session infrastructure is running
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
apiVersion: cluster.t4.dev/v1alpha1
kind: T4Workspace
metadata:
name: legacy-workspace
finalizers:
- cluster.t4.dev/workspace-protection
spec:
hostRef: legacy-host
displayName: Legacy workspace
owner: user:legacy-owner
repository:
repositoryId: legacy/project
ref: refs/heads/main
commit: 0123456789abcdef0123456789abcdef01234567
size: 20Gi
retentionPolicy: Retain
status:
observedGeneration: 4
pvcName: legacy-workspace-data
pvcPhase: Bound
capacity: 20Gi
phase: Ready
conditions:
- type: Ready
status: "True"
observedGeneration: 4
lastTransitionTime: "2026-01-15T12:00:00Z"
reason: StorageBound
message: Legacy workspace storage is ready
Loading
Loading