Order · stripe · rotation · offset · stale member
How a RAID array is rebuilt from member images. Every step, in the order the bench takes them, for anyone who has just watched a rebuild fail.
A failed RAID is put back together without writing a byte to any of its drives. Every member is imaged; the geometry is read from the metadata the controller or the software left on the members, or recovered from the parity and the data where it was cleared; the set is assembled in software from the images; and the file system is repaired on a copy of the virtual volume. This page says how each step is done and why, plainly, for anyone whose controller has just said Failed and who wants to know what is about to happen to the drives before deciding whether to send them.
Rather talk it through? An engineer answers the bench line
0800 6890668
Step one: every member is imaged, including the one that failed first.
Nothing is done to the set until every member has been copied. Each drive goes on the imager its interface needs, SAS, SATA, SCSI, Fibre Channel or NVMe, on hardware that controls how long a read may take and how many times it is retried, and is imaged sector by sector with its weak areas last and a map kept of every sector that could not be read. Drives with failed heads go to the clean bench first for a matched donor stack, a helium drive opened once and imaged immediately, a 520-byte drive imaged at its native sector size, a self-encrypting drive unlocked on the imager with the key you supplied. That is the drive site's work, and the same bench does it.
The drive that failed first is imaged too, and it matters more than people expect. When it dropped out, the set carried on writing without it, so its copy of the data is stale from that moment; but every stripe written before it dropped is still on it, exactly as it was, and a rebuild that stalled part-way never reached the stripes past the point it stopped. On a set that was degraded for a week before the second failure, the first-dropped drive holds most of the array unchanged. The map of unreadable sectors on each image is kept, because later it tells parity faults apart from media faults.
Step two: the metadata, where it survives.
Hardware controllers write the set's description to every member: PERC, MegaRAID and Intel in the DDF format near the end of the drive, HPE's Smart Array in its RAID Information Sector, Adaptec, Areca and Promise in formats of their own. mdadm keeps a superblock in one of four places depending on its version; ZFS keeps four labels per device with a ring of uberblocks in each; Storage Spaces keeps a pool database replicated on every disk; Windows dynamic disks keep the LDM database at the end of each. Where it survives, it gives the level, the member order, the stripe or chunk size, the parity rotation, the data offset, and on most formats an event count or sequence number for each member that says which are current and which dropped out first. The bench reads all of it from the images before assembling anything, and reads mdadm's superblock at all four locations, because an array re-created with a newer metadata version sometimes leaves the old superblock intact where the new one is not.
Step three: the geometry from the parity, where the metadata was cleared.
A configuration that was cleared, or overwritten by --create, or written by a controller the tools do not know, leaves the data and takes the description. The data is enough. On RAID 5, the parity block in every stripe is the XOR of the data blocks, so across all members at the same offset the blocks should XOR to zero on any stripe that was consistent when the set stopped; the bench tries candidate member orders and stripe sizes and the one that XORs to zero across the set is the right one. Where the parity block lands in successive stripes says whether the rotation is left or right, and whether the data restarts after the parity block (symmetric) or always from disk 0 (asymmetric). On RAID 6, P alone can be satisfied by a wrong layout by accident; Q, the second syndrome computed over a Galois field, cannot, and the geometry is believed only when both agree. The stripes where parity does not agree are the stale member's, or the write hole's, and the map of unreadable sectors says which.
Step four: the geometry from the data, where there is no parity.
RAID 0 and JBOD have nothing to XOR. The partition table and the file system's boot sector are on the first member, which finds it; the MFT on NTFS, the superblock and group descriptors on ext4, the VMFS headers, say where the next structures should be, and which member has them at which offset says the order and the stripe size. Entropy helps: at a stripe boundary, compressed video gives way to text or zeros, and the boundaries fall at multiples of the stripe. And the test is the largest files: a geometry that opens a multi-gigabyte file whole is right, because a wrong stripe size or order corrupts every file bigger than one stripe.
Step five: choosing the members.
With the geometry known, the bench chooses which image supplies each stripe. The current members supply everything they can read. The first-dropped member supplies only the stripes the survivors cannot give, and only after a check that the stripe was not written after it dropped, which the parity across the current members answers. A part-written replacement from a stalled rebuild supplies the stripes the rebuild did complete, which it computed correctly. On a mirror, the newest copy of each block is chosen from the event counts and the file system's own sequence numbers.
Step six: assembled in software, and the file system repaired on a copy.
The set is presented as one virtual volume built from the images, with nothing written to any of them. The file system is then checked and repaired on a copy of that volume: NTFS and ReFS journals replayed, ext4 and XFS metadata checked, Btrfs and ZFS walked from their newest consistent state with checksums honoured, VMFS opened and each virtual disk extracted with its snapshot chain, cluster shared volumes likewise with their checkpoint chains, SQL Server and Exchange databases recovered from their logs inside the extracted disks. Every repair is made on the copy; the images and the originals are not touched again. The list of what opened, and what did not, goes to you before any bill exists.
The tools.
The members are imaged on DeepSpar Disk Imager, Atola TaskForce and PC-3000 hardware. The geometry and the assembly use the tools the trade uses: R-Studio, UFS Explorer RAID Recovery, ReclaiMe, Runtime RAID Reconstructor, DiskInternals RAID Recovery, DMDE, Klennet ZFS Recovery and PC-3000's RAID module, chosen for the format in hand, and the bench's own scripts where a vendor's layout is not in any preset list. The file systems' own tools are used on copies of the virtual volume, never on the originals.
What makes a set unrecoverable.
More members lost than the level tolerates, with no usable stale member and no part-written replacement. A full or foreground initialisation, or a new array built and initialised over the old one. A rebuild or a consistency check that ran with the wrong geometry and rewrote parity from garbage. TRIM on SSD members after a deletion. Encryption with the key gone: SEDs without their controller's key, BitLocker without its recovery key. Each of those is found at the free look, said plainly, and on most jobs no data means no bill.
How long it takes.
Imaging scales with the members' size and health: healthy drives at their full speed, weak drives at whatever speed keeps them alive, a shelf of twenty-four over days. The geometry takes hours where the metadata survives and days where it must be recovered from parity across a large set. Assembly and file-system repair take a day or two, extraction of virtual machines and databases longer where there are many. A small set is usually 5–10 days at the bench after the free look, a shelf or a cluster 10–15 days at the bench, and the figure says which yours is.
The questions that come up first.
Is anything written to my drives?
No. Every member is imaged and every step after that is done on the images, with the file system repaired on a copy of the virtual volume. The originals go back to you untouched.
Why do you need the drive that failed first?
Because it holds every stripe written before it dropped, and a rebuild that stalled never reached the stripes past the point it stopped. It fills what the survivors cannot give.
Can you recover a set whose configuration was cleared?
Usually. The description is gone and the data is not; the geometry is recovered from the parity on RAID 5 and 6 and from the file system's own structures on RAID 0.
What if someone ran --create or chkdsk?
The bench works out what was written and where. --create rewrote the superblocks and possibly the data offset; chkdsk rewrote metadata around the wrong geometry. On images the rest is recoverable, and the file list says what those commands cost.
Does any of this cost me anything to find out?
No. The free look identifies the set, the geometry and what has been written since, and one figure follows in writing. £500 + VAT upwards for two to four members; from £1,250 + VAT for larger sets.
Now you know what the bench is about to do.
Send the form with the controller, the level, the members and what it says, and the first look tells you which of these your set needs, and what it would cost.