Skip to main content

VMware, Hyper-V and SAN data recovery for companies

Want to avoid making the situation worse?

Most important at the start

We recover data from VMware VMFS datastores, Hyper-V volumes, SAN LUNs and virtual-machine files when the underlying disks, RAID metadata or storage paths have failed.

Virtualisation failures are rarely a single-file problem. A missing datastore may involve controller metadata, a broken RAID set, unstable SAS/SATA media, a failed NAS head, a damaged VMDK/VHDX chain or stale snapshots after a power event. In business cases around Warsaw and across Poland, our role is to make the incident reversible before recovery work begins.

Therefore, the recovery of virtualization data requires a different approach than an isolated media incident. In this case, you need to understand the relationship between host, matrix, datastore, file system and virtual machine itself. The error on one layer easily moves higher.

First hour

Before changing the live environment, collect the facts and stop actions that rewrite storage metadata.

  • Save host models, matrix models and LUN layout,
  • Export or photograph controller, NAS, iSCSI, Fibre Channel and ESXi/Hyper-V error messages before they rotate out of logs.
  • Do not run any more rescans, import or VMFS/CSV modification procedures without making copies of metadata.

Most common crash scenarios

  • ESXi or Hyper-V stops seeing a datastore, CSV or LUN after restart.
  • A LUN is visible, but the VMFS volume cannot be mounted safely.
  • The array returns I/O errors, path flapping, APD/PDL events or disappearing devices.
  • A rescan, import or migration was started during an already unstable incident.
  • The hardware failure connected with a logical problem — e.g. a RAID error and damage to the datastore structure.

What to do in a productive environment

  1. Do not perform VMFS, CSV or datastore modification procedures without a full copy of the metadata.
  2. Don't map the LUNs again just to see if they come back.
  3. Do not run multiple changes at once: reboot hosts, rescans, imports, controller updates.
  4. Do not save new data to a resource that can be logically inconsistent.
  5. Do not assume that the problem is only on the hypervisor side if there have been matrix or power errors before.

What a safe recovery procedure looks like

Hypervisors and storage arrays are designed to keep services online, so they often keep retrying a failing path. That can be useful in normal operation, but dangerous during data loss. A forced rescan, rebuild, VMFS repair or LUN remount may overwrite metadata that explains how the environment was assembled.

We start by mapping the architecture: physical disks, RAID level, storage controller, SAN/NAS layer, hypervisor, datastore format and virtual-machine inventory. When media is unstable, each source disk or LUN is imaged first. Recovery then happens from verified copies, not from the only production evidence.

When to report the case right away

  • Do not run VMFS repair, CHKDSK or datastore repair tools on the only production copy.
  • the problem occurred after the blackout, matrix reboot or migration,
  • You see I/O errors, disappearing LUNs or read instability,
  • Someone has already tried to modify the production environment and symptoms change,
  • The stoppage grows and the environment has no current, independent copy.

Do you have a malfunction on the VMware, Hyper-V or SAN, and you don't want to take any chances on the production?
Give the server model, matrix, error messages and a description of the symptoms. On this basis, we will take a safe first step.

How to sort out the incident before the team loses control of changes in the production environment

In virtual environments, the biggest problem is rarely the failure itself. Much more often damage grows due to the pressure of time and parallel actions of several people: the administrator launches re-scanning of resources, the integrator re-maps LUN, another team launches the machine from copies, and users press to restore services. Therefore, at first it is worth appointing one person responsible for decisions, stopping uncontrolled changes and writing down a sequence of events.

It is also a good practice to separate priorities: which machines are critical, which data should be reproduced first, and which actions can wait until a safe diagnosis is made. Such order reduces the risk that action under time pressure will worsen the chances of full restoration of machinery or recovery of key data.

When Looking for External Help

If the matrix returned only partially, the datastore is visible, but inconsistent, and subsequent attempts to import or correct do not give a stable effect, then it is usually a sign that the problem does not end at host level. Then you need to analyze metadata, LUN order, controller cache status and relationship between the matrix layer and the VMFS or CSV file system. This is no longer a good time for uncontrolled operations performed on production.

If the subject also concerns the failure of the DKKNAS or the matrix in the company, see also manual first 24 hours after server/NAS failure and entry RAID failure at the company – what to do. These materials help to organize operations before the proper technical diagnosis begins.

Is it a problem with the layer of VMware / SAN or with the entire corporate infrastructure?

Send the symptoms, platform, datastore or LUN names and any logs you still have. We will help you decide what should be preserved before recovery work starts.

The most important pages in this cluster are listed below.

Before reporting a VM/SAN environment

Related guides about business storage failures

Before you submit a VMware, Hyper-V or SAN incident, compare it with the first-day server/NAS checklist and the common RAID backup misconception.

Server or NAS failure in a company - first 24 hours

How to stop chaos after a company storage incident and preserve facts before more repair attempts change the environment.

Read the guide

RAID is not a backup

Why redundancy does not replace a verified backup and how companies lose data when a rebuild becomes the plan.

Read the guide

Server or virtualisation incident? Start with controlled diagnostics.

Describe the affected server, array or virtualisation environment. The laboratory will indicate the safest diagnostic route, risk level and realistic recovery scope before work begins.

FAQ - VMware, Hyper-V and SAN data recovery

Do you need the whole storage system, or are exports and VM files enough?

Sometimes exports and VM files are enough, but after datastore, RAID, SAN or NAS failure we often need the original disks, LUN copies, controller information or logs. We confirm the safest handover path after the initial symptom review.

Can a VM be recovered after a host or datastore failure?

Often yes, if the datastore metadata, virtual disk chain or underlying storage can still be reconstructed from copies. The important part is to stop writing to the affected datastore before another rescan, import or repair attempt.

Do you work with VMDK, VHDX, snapshots and multi-machine environments?

Yes. We work with VMDK, VHDX, snapshot chains and multi-VM environments, including cases where the virtual machines depend on the same damaged datastore, RAID set or SAN/NAS layer.

How can I reduce risk after a virtual environment failure?

It is best to stop uncontrolled environmental modification activities, secure logs and describe the sequence of events. In VM environments, a mistaken sequence of actions can increase the extent of logical damage.