StorageGRID EC Rebalance fails immediately due to insufficient writable nodes
Applies to
- NetApp StorageGRID
- Erasure Coding (EC) Rebalance procedure
Issue
- The EC rebalance procedure fails immediately after starting new job or restating job with no moves attempted:
rebalance-data status========================JobID: 1234567890Site: SiteNameState: FailurePlanned Moves: 0Completed Moves: 0Failed Moves: 0Site Imbalance: UnknownRetry Rebalance: Yes========================
-
From Primary Admin node
/var/local/log/rebalance-data.logrecords No jobs found for EC repairs:D, [2026-04-29T12:31:32.611561 #4050370] DEBUG -- : No jobs found for replicated repair for the given parameters {"alert_for_node_repair"=>false, "include_node_repair"=>true, "repair"=>["replicated", "erasure-coded"], "state"=>"running"}D, [2026-04-29T12:31:32.634167 #4050370] DEBUG -- : No jobs found for erasure-coded repair for the given parameters {"alert_for_node_repair"=>false, "include_node_repair"=>true, "repair"=>["replicated", "erasure-coded"], "state"=>"running"}I, [2026-04-29T12:31:32.723691 #4050370] INFO -- : EC Site rebalance is restarted successfully.
From EC Job Leader storage node
/var/local/log/bycast.logrecords Exhausted retry limit:Apr 29 12:28:44 nodename ADE: |21014662 1884999497 ECJM ???? 2026-04-29T12:28:44.205607| NOTICE 0040 ECJM: Starting job 1234567890: Site Rebalance - Group ID 20.Apr 29 12:28:44 nodename ADE: |21014662 1884999497 ECJM ^RDY 2026-04-29T12:28:44.231500| NOTICE 0234 ECJM: 1234567890(rebalance 20): RetryingApr 29 12:28:44 nodename ADE: |21014662 1884999497 ECJM ^RDY 2026-04-29T12:28:44.231579| NOTICE 1283 ECJM: 1234567890(rebalance 20): saving state. status: JOBSTATUS_IN_PROGRESSApr 29 12:28:44 nodename ADE: |21014662 1884999497 ECJM EVIU 2026-04-29T12:28:44.254866| NOTICE 0281 ECJM: 1234567890(rebalance 20): Already have 0 recommendations, asking for a further 200...Apr 29 12:28:46 nodename ADE: |21014662 1884999497 ECJM EVIU 2026-04-29T12:28:46.179064| NOTICE 0295 ECJM: 1234567890(rebalance 20): received message GVCRApr 29 12:28:46 nodename ADE: |21014662 1884999497 ECJM EVIU 2026-04-29T12:28:46.179083| WARNING 0300 ECJM: 1234567890(rebalance 20): Got error SUNV while trying to retrieve VCS moves. Retrying.Apr 29 12:28:46 nodename ADE: |21014662 1884999497 ECJM GVCR 2026-04-29T12:28:46.179112| NOTICE 1283 ECJM: 1234567890(rebalance 20): saving state. status: JOBSTATUS_PAUSEDApr 29 12:28:46 nodename ADE: |21014662 1884999497 ECJM ^RDY 2026-04-29T12:28:46.270051| WARNING 0062 ECJM: Caught exception 'ENFORCE failed: !"Exhausted retry limit or retry time for getting move recommendations"' when running job 1234567890: Site Rebalance - Group ID 20.Apr 29 12:28:46 nodename ADE: |21014662 1884999497 ECJM ^RDY 2026-04-29T12:28:46.292765| ERROR 1125 PROC: Exception: /build/src/modules/ErasureCoding/EC_JobManager_Module/SiteRebalanceJob.cc(382): Throw in function getMoveRecommendations#012Dynamic exception type: boost::wrapexcept<std::runtime_error>#012std::exception::what: ENFORCE failed: !"Exhausted retry limit or retry time for getting move recommendations"#012- Some storage nodes are in a Read-Only state.
ILM placement unachievablealerts may also be present.
