Emergency server help: get in touch

ESXi APD PDL: Datastore Not Accessible, Safe Recovery Steps

Diagnose and recover from ESXi datastore inaccessibility caused by All Paths Down or Permanent Device Loss on ESXi 7, 8 and 9, using esxcli storage commands, vmkernel.log, path rescans, APD timeout settings and vSphere HA component protection.

Published Updated 6 min read

When a datastore turns grey in the vSphere Client and VMs on it freeze or crash, ESXi has lost access to the underlying LUN or NFS export. The host classifies the loss as one of two states. All Paths Down (APD) means every path to the device stopped responding without the array saying why, so the host keeps retrying because the storage might return. Permanent Device Loss (PDL) means the array answered with a SCSI sense code saying the LUN is gone, so the host stops trying and marks the device dead. Recovery differs. This guide covers ESXi 7.0, 8.0 U3 and 9.x with Fibre Channel, iSCSI and NFS storage.

Applies to ESXi 7.0, 8.0 U3 and 9.x with Fibre Channel, iSCSI and NFS storage

Short answer: Check /var/log/vmkernel.log and esxcli storage core device list to determine whether the device is in APD (“APD Timeout” messages, status “Dead Timeout”) or PDL (“Permanent Device Loss” messages, status “Not Consumed” or “Dead”). For APD, fix the fabric, switch, target or network issue and run esxcli storage core adapter rescan --all; the host reconnects automatically once paths return. For PDL, the LUN was intentionally removed or unmapped; power off or migrate any VMs still referencing it, unmount the datastore, detach the device with esxcli storage core device set --state=off -d <naa.id>, and rescan so the host clears the stale entry.

Tell APD from PDL

Search the vmkernel log on the affected host:

grep -i "APD\|PDL\|Permanent Device Loss\|all paths down" /var/log/vmkernel.log | tail -40
esxcli storage core device list | grep -B2 -A10 "naa.6001"
esxcli storage core path list -d naa.6001405a1b2c3d4e
esxcli storage nfs list

In the device list, Status: Dead Timeout or Is APD: true-style text indicates APD; a PDL device shows Status: Not Consumed with a note that it is in a permanent device loss state. In the path list, APD shows every path as dead; PDL shows paths active but the device reporting an unrecoverable sense code (0x5 0x25 00, “logical unit not supported”). NFS datastores only ever go APD because NFS has no equivalent sense code.

Recover from APD

APD is almost always external: a switch reboot, a zoning or VLAN change, an iSCSI target restart, an array controller failover that took too long, or a NIC failure on an iSCSI vmkernel port. Confirm connectivity from the host:

vmkping -I vmk1 192.168.50.10
esxcli iscsi session list
esxcli storage san fc list
esxcli storage core adapter rescan --all
esxcli storage filesystem list

Once paths return, the device leaves APD without intervention and VMs that were blocked on I/O resume. VMs whose guests timed out on disk I/O may have read-only file systems or Windows crashes and need a guest reboot.

The host’s behaviour during APD is governed by the advanced settings Misc.APDHandlingEnable (1) and Misc.APDTimeout (140 seconds by default). After the timeout, non-VM I/O is failed fast so hostd does not hang; VM I/O keeps retrying. Do not disable APD handling to “make VMs survive longer”; it just makes hostd unresponsive.

Recover from PDL

PDL follows an intentional removal that skipped the unmount step, or an array-side failure. If VMs are still registered on the datastore, they need to be powered off (they cannot be migrated because their disks are gone):

esxcli storage filesystem list
vim-cmd vmsvc/getallvms | grep "datastore-name"
esxcli vm process list
esxcli vm process kill --type=force --world-id=<id>

Then unmount and detach:

esxcli storage filesystem unmount -l DatastoreName
esxcli storage core device set --state=off -d naa.6001405a1b2c3d4e
esxcli storage core adapter rescan --all
esxcli storage core device detached list
esxcli storage core device detached remove -d naa.6001405a1b2c3d4e

The detached list step matters: ESXi remembers detached devices so a reappearing LUN is not automatically reclaimed, which confuses people when the same LUN is re-presented later.

Let HA respond automatically

vSphere HA can restart VMs elsewhere when a host loses storage. In the cluster » Configure » vSphere Availability » Failures and responses » Datastore with PDL, choose “Power off and restart VMs”; for Datastore with APD, choose “Power off and restart VMs (conservative)” with a response delay of a few minutes. This VM Component Protection keeps a fabric blip from triggering restarts while ensuring a true loss is handled without a human.

Verify

After recovery, esxcli storage filesystem list should show the datastore mounted with Mounted: true, esxcli storage core path list should show all expected paths active, and the datastore should be green in the client. For PDL clean-up, the device should no longer appear in the device list at all. Rerun the vmkernel grep to confirm no new APD or PDL lines appear over the following minutes.

Common pitfall

Rebooting the host during APD is the reflex to resist. The host will hang at boot waiting for the missing devices, and in the meantime VMs that would have resumed when the paths returned are down for the duration. Fix the storage path first. Equally, treating a PDL as an APD and waiting for the LUN to return leaves zombie VMs on the host until someone kills them; see Kill an unresponsive VM on ESXi.

ESXi APD PDL at a glance

ESXi APD PDL summary card: Check /var/log/vmkernel.log and esxcli storage core device list to determine whether the device is in APD ("APD…
In short: Check /var/log/vmkernel.log and esxcli storage core device list to determine whether the device is in APD (“APD Timeout” messages, status “Dead Timeout”) or PDL (“Permanent Device Loss” messages, status “Not Consumed” or “Dead”).

Official documentation: Broadcom TechDocs (VMware), Linux man pages.

Related guides: Fix “Another task is already in progress” in vCenter with vim-cmd task_info · Kill an unresponsive VM on ESXi with vim-cmd, esxcli and esxtop · Fix “The redo log of .vmdk is corrupt” on ESXi and Horizon linked clones.

Frequently asked questions

Does APD also affect vSAN datastores?

vSAN handles component loss through its own resync and absent/degraded states rather than the APD and PDL mechanism. A vSAN host that loses the vSAN network shows objects as inaccessible in the Skyline health view, and the recovery is to restore the network rather than to rescan adapters.

How long does ESXi wait before declaring APD timeout?

140 seconds by default (Misc.APDTimeout), after which non-VM I/O fails fast. VM I/O continues to retry indefinitely, so guests may hang until the storage returns or HA intervenes.

Can I undo detaching a device after a PDL clean-up?

Yes. esxcli storage core device set --state=on -d <naa.id> reattaches a detached device, and removing it from the detached list with esxcli storage core device detached remove lets a rescan claim it again if the LUN is presented.

Maintenance record

This guide changes servers, data or security settings, so we re-check it against current versions on a fixed schedule. Take a backup or snapshot before you start.

Maintained by
srvScripts editorial team
Supported versions
ESXi 7.0, 8.0 U3 and 9.x with Fibre Channel, iSCSI and NFS storage
Last full review
Next review

Free website test

Is your website set up right?

Check SSL, security headers, redirects, robots.txt, sitemap, llms.txt and security.txt in one test. It takes about 30 seconds.