Skip to main content

Upgrade and Downgrade a Cluster

PhoenixAI requires its components to change version in a fixed order:

  • Upgrade: CN first, then FE. CN is backward compatible with an older FE; an FE that is newer than its CN is not supported.
  • Downgrade: FE first, then CN. The mirror image of the rule above.

The operator enforces this order for you. You change the image versions in the PhoenixAICluster (and in any PhoenixAIWarehouse or PhoenixAICnGroup) in one apply, and the operator rolls the components in the right sequence, waiting for one to finish before it starts the other.

What the operator does​

Before it writes a StatefulSet whose image version changed, the operator asks FE which version every node is actually running (SHOW FRONTENDS and SHOW COMPUTE NODES) and checks one rule:

Component being writtenRule
FEthe new FE version must not be higher than any running compute node
CN (cluster, warehouse or CN group)the new CN version must not be lower than any running FE

If the rule does not hold yet, the write is postponed and retried on the next reconcile. Nothing else in the reconcile is skipped: services, configuration and autoscaling keep being reconciled.

Put together, the two rules produce the required order without you doing anything:

  • Upgrade FE and CN together: CN rolls first. FE waits until every compute node reports the new version, then rolls.
  • Downgrade FE and CN together: FE rolls first. CN waits until every FE reports the old version.
  • Resume an interrupted upgrade (CN already new, FE still old): FE rolls right away.
  • Roll back the FE change while CN is still rolling: allowed at once; nothing to wait for.

The rule is also applied to the spec itself, before anything runs: the FE image version must not be higher than the CN image version of the cluster, of any warehouse attached to it, or of any CN group attached to it. This includes CN groups scaled to 0 replicas. A parked group with an old image would register old nodes into a new FE the moment you scale it out, so raise its image together with the others (this rolls nothing, there are no pods).

Only the version in the image tag matters (4.1.16, 4.1.16-ee, 4.1.16-ee-4a26e40 are all version 4.1.16). Changes that do not change the version, such as environment variables, resources or configuration, roll FE and CN in parallel exactly as before. Scaling a CN StatefulSet in or out, by hand or through an autoscaler, is never held back.

What you see while it waits​

The wait is reported in the component status of the object whose write is postponed, under a RolloutOrderBlocked: prefix, and the component phase is reconciling:

kubectl get pac <cluster> -n <namespace> -o jsonpath='{.status.phoenixAIFeStatus.reason}{"\n"}'
# RolloutOrderBlocked: FE target 4.2.0 must not be higher than any running compute node; waiting for cn-1.…=4.1.16

kubectl get paw <warehouse> -n <namespace> -o jsonpath='{.status.reason}{"\n"}'
kubectl get pacg <cngroup> -n <namespace> -o jsonpath='{.status.reason}{"\n"}'

A spec that breaks the rule is reported the same way with an InvalidImageOrder: prefix and names the compute owner whose image is too low:

InvalidImageOrder: FE target 4.2.0 is higher than the image of PhoenixAICnGroup az-a=4.1.16; FE must never be newer than any compute node, raise those images first (an owner with 0 replicas must be raised too)

The operator does not emit Kubernetes events for either state. Waiting is the normal, temporary condition of a combined upgrade, not an error. The reason disappears as soon as the write goes through.

When the operator cannot ask FE​

The check needs a working SQL connection to FE, using the root password the cluster was deployed with (the MYSQL_PWD environment variable the Helm chart injects when initPassword is enabled). When FE cannot be reached, or the password is rejected, the operator cannot prove the rule and therefore postpones the write, saying so in the reason:

RolloutOrderBlocked: cannot verify runtime versions (dial tcp …: connection refused); fix FE first, or set annotation starrocks.com/skip-rollout-order-check="true" on PhoenixAICluster <name> to override
RolloutOrderBlocked: authentication to FE failed (Error 1045 …); check the root password secret, or set annotation …

If you changed the root password, update the secret the cluster references first. If FE is down and the fix is itself an FE image change (for example, rolling back an FE upgrade that failed to start), use the override below.

Overriding the check​

Add this annotation to the object whose StatefulSet write you want to let through unchecked:

kubectl annotate pac <cluster> -n <namespace> starrocks.com/skip-rollout-order-check="true"

It works on PhoenixAICluster, PhoenixAIWarehouse and PhoenixAICnGroup, and only affects that object's own writes. Remove it once the cluster is healthy again:

kubectl annotate pac <cluster> -n <namespace> starrocks.com/skip-rollout-order-check-

There is no operator-wide switch: the annotation is deliberately the only way to bypass the check.

Notes​

  • Image tags without a version (latest, a digest, a private build name) cannot be compared, so the check does not apply to them and the write proceeds as before. The operator logs rollout order check skipped: image tag not parseable. Use versioned tags to get ordered rollouts.
  • Hand-added compute nodes count too: every node FE lists constrains the FE version, whether the operator manages it or not.
  • spec.waitForFullRollout keeps its existing meaning: when set, BE/CN changes wait for the FE StatefulSet to finish rolling, for every kind of change. It is independent of the version ordering described here, and it now holds in the very reconcile that changes FE as well (it used to let a CN change through in that one pass).
  • Load on FE: the operator queries FE only while a version change is pending, at most once every 10 seconds per cluster, and marks its statements with /* phoenixai-operator: rollout-order-check */ so they can be filtered in the FE audit log.