asked on

ESXi 4.1 post upgrade disk latency issue

Upgraded our production server from ESXi 4.0 to 4.1 over the weekend. Since the upgrade I have been seeing an issue where disk latency to our iSCSI SAN will intermittently spike dramatically for about 40 seconds. When this happens all other VM performance metrics flatline to 0.

As our SBS 2008 VM is one of those affected this means the network just stops and then starts up again. At the moment the issue is just annoying, I'd like to get it sorted before it escalates.

This issue did not exist while we were running ESXi 4.0 on this host. Any pointers on where to look at resolving this issue are much appreciated.

ASKER CERTIFIED SOLUTION

Danny McDaniel

membership

This solution is only available to members.

To access this solution, you must be a member of Experts Exchange.

Start Free Trial

SOLUTION

mastoo

membership

This solution is only available to members.

To access this solution, you must be a member of Experts Exchange.

Start Free Trial

siht

ASKER

I could vmkping the iscsi ip from the console during the spikes and there was nothing jumping out at me from the logs.
There was a bit of this:

[2011-03-23 10:27:12.729 1BBEEB90 verbose 'App'] [VpxaMoVm::CheckMoVm] did not find a VM with ID 5 in the vmList
[2011-03-23 10:27:12.729 1BBEEB90 verbose 'App'] [VpxaAlarm] VM with vmid = 5 not found
[2011-03-23 10:27:12.729 1BBEEB90 verbose 'App'] [VpxaMoVm::CheckMoVm] did not find a VM with ID 8 in the vmList
[2011-03-23 10:27:12.729 1BBEEB90 verbose 'App'] [VpxaAlarm] VM with vmid = 8 not found
[2011-03-23 10:27:12.729 1BBEEB90 verbose 'App'] [VpxaMoVm::CheckMoVm] did not find a VM with ID 9 in the vmList
[2011-03-23 10:27:12.729 1BBEEB90 verbose 'App'] [VpxaAlarm] VM with vmid = 9 not found
[2011-03-23 10:27:22.727 1BBEEB90 verbose 'App'] [VpxaMoVm::CheckMoVm] did not find a VM with ID 5 in the vmList
[2011-03-23 10:27:22.727 1BBEEB90 verbose 'App'] [VpxaAlarm] VM with vmid = 5 not found
[2011-03-23 10:27:22.727 1BBEEB90 verbose 'App'] [VpxaMoVm::CheckMoVm] did not find a VM with ID 8 in the vmList
[2011-03-23 10:27:22.727 1BBEEB90 verbose 'App'] [VpxaAlarm] VM with vmid = 8 not found
[2011-03-23 10:27:22.727 1BBEEB90 verbose 'App'] [VpxaMoVm::CheckMoVm] did not find a VM with ID 9 in the vmList
[2011-03-23 10:27:22.727 1BBEEB90 verbose 'App'] [VpxaAlarm] VM with vmid = 9 not found

But it was not especially associated with the latency events.
I have migrated the SBS VM off of the iSCSI storage and restarted the storage. The errors above have gone away and while the latency on the iSCSI datastore still spikes it never hits higher than 12 ms, as opposed to the 15000 ms spikes I was seeing.

I'll migrate the VM's back to iSCSI slowly & keep an eye on it. This issue occurred under standard load, we'd been running those VM's on the iSCSI LUN for months. I'd like to have gotten to the bottom of it but I'm glad it's gone away for now.

Danny McDaniel

-did you patch the VM recently?

-is it running on a snapshot?

-you might want to setup a syslog server to get better logging. you can use the VMA appliance to do that job if you want.

siht

ASKER

-did you patch the VM recently?

No and this was affecting all VM's on the iSCSI storage, the SBS VM was the biggest issue. The issue was definitely triggered by upgrading the host to ESXi 4.1.

-is it running on a snapshot?

No snaps on production VM's.

-you might want to setup a syslog server to get better logging. you can use the VMA appliance to do that job if you want.

Way ahead of you there :)

siht

ASKER

Issue was not resolved but answers were still helpful.