Skip to main content

VMware ESXi cannot see the datastore after restart: LUN recovery case

VMware ESXi datastore missing after restart and LUN recovery

Restarting the VMware server after a power failure can end in a lack of access to datastore and a whole group of virtual machines. This is a critical scenario, because one hastily action after a take-off — re-scan, rebuilding the RAID or writing in the wrong place — could make the metadata worse.

The cause may be damaged VMFS metadata, a multipath issue, stale signatures, a RAID controller fault, inconsistent SAN state or corruption of the partition table that points ESXi to the datastore.

The LUN can appear online on the array while vSphere still shows the datastore as unknown, unmounted or missing. Visibility of a LUN is not proof that the VMFS volume is safe to modify. B2B paths for companies and the analysis of RAID/NAS, without recovery of RAID before confirmation of array and LUN state.

Why a datastore disappears after an ESXi restart

After a power failure, ESXi can see LUN as available, but not mount datastore VMFS. The cause is damage to VMFS metadata, multiplepath path error, problem with the RAID controller or inconsistencies on the side of the SAN array.

In a typical incident, power returns, the host starts, but the datastore does not mount and dozens of virtual machines remain offline. The temptation to rescan, rebuild paths or recreate inventory is high.

Symptoms of datastore loss after a restart

The better route is staged and reversible: preserve disk and LUN state, reconstruct the storage layout if needed, identify VMFS metadata damage and verify virtual machines from recovered copies.

Do not delete and recreate the datastore, format VMFS, rebuild RAID blindly or move VMDK files before the metadata state is understood. Do not keep rescanning paths if each attempt changes logs and state without bringing the datastore back.

Step by step: recovering a VMware datastore

After data is secured, plan infrastructure repair separately from recovery. The goal is to restore virtual machines from verified copies and avoid returning a damaged storage layer to production.

Diagnosis includes disk imaging, analysis of the structure of the RAID and reproduction of the array system in the working environment. Only on such a prepared material can you look for the beginning of the VMFS partition and assess whether you can safely restore access to virtual machines.

In such cases, the order of action is important. The aim is not to quickly try to repair on production, but to keep the environment as fully as possible, reduce the risk of overwriting structures and work on a working copy.

Case study: 40 virtual machines after an ESXi restart

In similar failures after the planned reboot of the VMware server, some datastores may disappear and virtual machines do not start despite the visible LUN. An earlier rescan storage adapter or an attempt to re-create a datastore may breach metadata, so further action should begin by securing the state and reconstructing the system outside the production environment.

In the laboratory approach, the process is divided into stages: media security, imaging, RAID or LUN reconstruction, VMFS analysis and machine verification. It is only after such a diagnosis that it is possible to determine which range of data is truly available and how to safely plan the further launch of services.

What to prepare before contacting the lab

For corporate environments the most important are: array topology, host count, disk order, recent administrative operations and the exact moment of loss of datastore. It is important to collect logs, discharges of messages from ESXi, and to know whether rescan, remount or any attempts to repair VMFS were performed after the reboot. This makes it easier to determine whether the problem concerns the logical layer, array or LUN itself. Entries about recovery of data from VMware, Hyper-V and SAN and RAID failure at the company.

When not to improvise after a host restart

If datastore continues to disappear, some virtual machines do not see files, and the storage panel shows errors or degraded condition, each subsequent reboot and any attempt to repair "on production"increases the risk of downtime. In such situations, it is better to suspend further operations, secure configuration information and move to orderly diagnostics. When the environment supports accounting systems, ERPs or backups, quick contact and a description of the host, datastore and LUN before the problem goes from logical to multilayer infrastructure failure.

How to close the incident without adding business risk

If the VMware environment supports production, accounting system, ERP or key virtual machines, not only recovery itself, but also the pace and order of operations. It is best to choose an orderly path immediately:

It's safer than another reboot, rescan, or improvised reconstruction in a working production environment.

Datastore, SAN or virtual machine recovery?

This case concerns the environment of virtualization and LUNs on the corporate side. If the datastore has disappeared after the reboot or the storage layer is unstable, go to the service path for the environments of VMware, SAN and servers.

The most important pages in this cluster are listed below.

Datastore missing after restart?

Describe host, array, LUN, datastore, ESXi logs and whether the rescan, remount, or restoration of the RAID has been performed. Do not run any more repairs on production; we are working on a copy, and the technician will indicate a safe sequence of actions.

Talk business