Company RAID failure: first response before rebuild
The worst RAID decisions usually happen in a hurry: a Warsaw office cannot open invoices, a NAS starts beeping, VMware loses a datastore, and somebody clicks Rebuild because the panel offers it. Before that click, stop and record the array state. Disk order and metadata may be more valuable than another restart.
First rule after a RAID alert
Hold virtual machines, backups and all records, list the RAID level, device model, disk order and exact error message. This is usually more important than rebooting under time pressure or rebuilding the RAID performed on the original array.
What not to do in the first minutes
- Do not run RAID reconstruction, Resync or re-initialization of the matrix.
- Do not move disks between bays to “test” whether the NAS sees them.
- Do not replace a second drive because the first rebuild failed.
- Do not accept controller or NAS prompts if you do not know what metadata they will write.
What to record before anything changes
- First, stop services, virtual machines, backups and all processes that record data on the matrix.
- Save controller model or NAS, RAID level and current error messages.
- Mark disks: sinus order, serial numbers, position in the housing.
- Don't start the reconstruction on the originals until you know what really went wrong.
Most common crash scenarios
- after one drive drops out, but another disk already has pending sectors.
- Two disks began reporting bugs and the matrix stopped folding logically.
- The controller or NAS overwritten or lost metadata after reboot or upgrade.
- The administrator has launched a RAID reconstruction on the wrong disk or after a bad diagnosis.
Why the Restoration of RAID does not always help
Take photos of the bay order, drive labels and screen messages. Export logs if the NAS or controller allows it without writing to the array. Record the RAID level, number of disks, device model, serial numbers, volume names and the time when symptoms started.
When the case should go to the lab
For company environments, also list the services that used the array: accounting, SQL databases, ERP, file shares, surveillance recordings, VMware, Hyper-V, backup repositories or user profiles. This helps set priorities before diagnostics begin.
Is RAID or NAS off-line and the company stands?
Describe device model, RAID level, disk count and exact error message. This allows you to quickly assess whether the problem concerns one disk, matrix metadata or controller.
Safe first steps before contacting the lab
A rebuild is designed for a known single-disk failure with the remaining drives healthy. Real incidents are often messier. RAID 5 may have a second disk with unreadable sectors. RAID 6 may have stale metadata after a power event. A NAS may mark the wrong disk as failed after a controller or firmware problem.
When the failure of the RAID is more complex than it looks
If the rebuild reads unstable disks for hours, it can stress the remaining members and write a new, incorrect state. In recovery work, an untouched inconsistent array is usually easier to analyse than an array that has been rebuilt several times in the wrong order.
How to brief management and users
Do not rely only on the colour of the LED or a single dashboard label. Export or photograph the storage pool status, RAID group state, disk serial numbers and recent system events if the panel is still responsive. If the NAS suggests repair, migration or expansion, pause until you know whether it will write new metadata.
When Stopping Does Not Explain Uncontrolled Actions
If production, sales system, accounting system or company backup work on the matrix, stop improvisation. It is better to collect a set of information and enter the path of diagnosis than to run another reconstruction without disc image and order control.
When the failed array is also the backup repository
- RAID 5 shows degraded even though the drives look healthy
- Ransomware on NAS/QNAP: how not to make the situation worse
- VMware ESXi cannot see the datastore after restart
What else to collect before diagnosis of RAID
Label every drive before it leaves the enclosure: bay number, serial number and original position. Pack disks separately so labels remain readable. In RAID recovery, a correct disk-order photo can save hours of reverse engineering and may prevent a wrong reconstruction attempt.
In the laboratory, it is also important whether the failure was a one-time one or a growing one: the single disk has been dropping for weeks, the RAID has been operating long in degraded mode, someone has replaced the storage device with a larger one or the controller has started rebuilding itself. Such a story helps distinguish between the simple lack of one disk and incorrect reconstruction, damaged metadata or the problem of several media at once.
Promise limit in company failure
The lab does not need the NAS panel to “look healthy”. It needs stable images of member disks, metadata fragments, parity layout, stripe size and the sequence of events. That is why the original state matters more than a rushed attempt to make the enclosure mount again.
What to send with the case description
With RAID, the order of the media is part of the data. Before anyone takes the disks out of their pocket, mark the bays and serial numbers: a photo, a label or a simple table. Do not assume that the NAS panel always shows the correct order after restart or migration of the controller.
If the array holds accounting data, production files, SQL databases, virtual machines or a whole-company backup, do not treat it as a casual repair. A laboratory workflow starts by imaging member disks and reconstructing the array from copies, not by experimenting on the originals.
It is also worth determining one person on the side of the company who collects information and makes decisions. When the RAID fails, the lack of one decision path can be as dangerous as a technical error: one administrator restarts, the other replaces the disk, the third triggers the backup task. One decision path limits random changes.
Send the RAID level, number of drives, device model, disk order, error messages, last known good state and actions already taken. If the array contains personal data or business-critical files, describe the priority folders and any legal or operational deadline. Clear context helps avoid diagnosis based on guesswork.