Tonight, in this order
The array has failed. Here is the first response. Five things to do, in order, and eight things not to, each with its reason.
Every message a controller or a software stack gives when an array fails comes with a button that makes it worse, and the buttons are the same across every vendor's dialect: rebuild, force online, import, clear, initialise, repair. This page is the first response in one place, in the order the bench would give it on the telephone at three in the morning, followed by the wrong responses and the reason each one costs. Nothing on it needs a tool, a download or a command; it is about not writing to the drives until there is a copy.
Rather talk it through? An engineer answers the bench line
0800 6890668
The five things to do, in order.
Open a case →Stop writes, and do not reboot repeatedly Free
If the set is degraded and still serving, copy what matters most to somewhere else, now, and then stop writing to it. Do not reboot to see whether the message clears; each boot is a chance for the controller to start a rebuild, and each spin-up of a weak drive costs it sectors. If the set is offline, leave it offline.
Press nothing the controller offers
Not Rebuild, not Force online, not Import, not Clear, not F2 at an HP 1779 prompt, not Repair on a NAS, not Discard preserved cache. Every one of them writes to the drives or throws something away. Let a POST prompt time out; the default is the safe key.
Photograph the screen and label the slots
Photograph the controller's screen with every message and bay number it shows. Before any drive comes out, write its slot number on it with a marker, and keep it in its carrier. The metadata usually knows the order; the label makes it certain, and on RAID 0 it is the only record.
Export the log without changing state
PERC: the TTYLOG from OpenManage or iDRAC, or perccli /c0 show all. MegaRAID: storcli /c0 show all and the event log. HPE: an SSA diagnostic report or the iLO log. mdadm: --examine on every member, /proc/mdstat and dmesg. ZFS: zpool status -v and zpool import with no flags. Storage Spaces: Get-StoragePool, Get-VirtualDisk, Get-PhysicalDisk. vSAN: the health report and the object list. All of them are read-only; none of them assembles, imports or repairs anything.
Power down, and send every member
Every drive in the set, including the one that failed first and any part-written replacement, each in an anti-static bag in its own padding; the controller only if it was replaced and the set is now foreign, or if it holds an encryption key or preserved cache. The posting address comes by email in reply to the form. The first look is free.
The eight wrong responses, and what each one costs.
| The response | What it does to the drives | Why it costs |
|---|---|---|
| Rebuild, resync, resilver, Repair | Reads every sector of every survivor; writes to the replacement | On drives bought together, a survivor's read error is where it stops; the set goes from degraded to failed |
| Force a drive online | Serves a drive that could not answer, or is stale, as current | Weeks-old stripes mixed into current data; the file system writes on top |
| Clear the foreign configuration | Deletes the set's description from every member | The geometry becomes a reconstruction; the next step, a new virtual disk, initialises |
| Import with a member missing | Brings up a degraded set the controller then wants to rebuild | As for rebuild, with the stale member possibly included |
| Delete and recreate, then initialise | Rewrites the metadata; a quick init writes the start and end; a full or background init rewrites the set | The start of the volume at least, and everything if the init runs |
| chkdsk or fsck on a degraded or mis-assembled volume | Repairs what it finds and writes the repairs back | On the wrong geometry or with a member missing, it repairs the wrong thing onto the right data; Microsoft's 2020 guidance on RAW parity spaces was not to run it |
| zpool import -F, -X or -T on the originals | Rewinds the pool by discarding the newest transaction groups | On the originals, those groups are gone; on images, the same rewind is reversible |
| mdadm --create --assume-clean, or --assemble --force | Writes new superblocks with typed geometry and possibly a new data offset; or promotes a stale member | The description no longer matches the data; or old data is served as current |
Why the first response is the same whatever the controller says.
A degraded set is serving from its survivors with no margin. A failed set has stopped serving, and the data on its members is exactly what it was when the last tolerable member dropped. Neither state is a loss. What every wrong response above has in common is that it writes: to the replacement, to the survivors, to the metadata, to the file system, or to the pool's history. On images, every one of those writes would be reversible; on the originals, none of them is. The bench's entire method is to make the copy first and do everything on it, and the first response is simply the part of that method you can do tonight: stop the writes, keep the record, and get every member to the imager unchanged.
The vendors' own words are on the message pages: HPE's F2 to accept data loss, Dell's rebuild with errors, Synology's can no longer repair it by yourself, QNAP's does not help in the event of disk failure, Microsoft's lost data because too many drives failed. Each is a description of what the vendor's tool can do, not of what is on the drives. How a RAID array is rebuilt from member images says what the bench does instead, step by step.
The questions that come up first.
The set is degraded but still working. Should I rebuild?
Only if every surviving member is healthy and you have a backup. If a second drive shows predictive failure, the rebuild reads the bad sectors that finish it. Copy what matters, power down, and image the weak drives first.
Is it safe to export the log?
Yes. Every command listed above is read-only: it reports the controller's or the software's view without assembling, importing or repairing anything. Copy the output and send it with the form.
What if I already pressed one of the buttons?
Stop now, and write down what was pressed and in what order. On images, most of what it did is recoverable; the bench works out what it reached. The worst outcome is pressing the next one.
Do I send the controller?
Only if it was replaced and the set is now foreign, or if it holds an encryption key or preserved cache. The set's description is on the drives.
What does it cost?
A set of two to four members is £500 + VAT upwards after the free look, fixed in writing; larger sets, 15K SAS sets, storage shelves and enterprise storage from £1,250 + VAT. On most jobs no data means no bill.
Nothing gets worse while it is powered down.
Stop the writes, keep the record, and send every member. The first look is free, and the figure is in writing before anything chargeable happens.