Virtual Machines Enter Paused State on OpenStack Host After ONTAP Upgrade
Applies to
- NetApp AFF-A700
- ONTAP 9.15.1P8 (Cluster-Mode)
- OpenStack/KVM environments using Platform9 VM-HA
- iSCSI-attached Linux compute nodes
Issue
After performing an ONTAP upgrade, a single OpenStack/KVM compute host experienced an outage. All virtual machines running on this host entered a paused state. The issue was resolved after the host was rebooted.
Relevant log output from /var/log/pf9/pf9-vmha-agent.log:
ERROR "unable to send payload to hamgr" status=503 INFO "Not able to reach hamgr" INFO "All hosts and DU not reachable assuming current host is offline, pausing all domains" INFO Running command "[sudo virsh suspend <uuid>]" (repeated for all 28 domains)
Additional log output from /var/log/pf9/ostackhost.log:
ERROR nova.virt.driver [...] Exception dispatching event <LifecycleEvent:...Paused>: Remote error: DBConnectionError (pymysql.err.OperationalError) (2013, 'Lost connection to MySQL server during query')
