Taking work now — the first look is freeWhole sets posted in from anywhere in the UK, or handed in at ten drop-off pointsQuicker still, give us a ring:0800 6890668
RARRAID Array Data Recovery 0800 6890668 Price my job
RAR / Whatever it is saying now / Rebuild failed: a second drive went during it

rebuild failed · rebuild stalled at n% · second drive failed · URE · 10^14 · 10^15 · hot spare · punctured

The rebuild failed, because a second drive went during it. The commonest way an array arrives here, and the one with the most left to recover.

A rebuild computes the missing member from every survivor, stripe by stripe, reading every sector of every one of them and writing to the replacement. On a set whose drives were bought together, it is the heaviest read those drives will ever do, at the moment one of them has already failed. A survivor that returns an unrecoverable read error, or that has been predictive for months, is where it stops: the rebuild stalls at a percentage, the controller drops the survivor, and the virtual disk goes from degraded to failed. The maths is honest about why. A drive rated at one unrecoverable error per 10^14 bits is rated to meet one in about 12.5TB read, and a rebuild of four 8TB drives reads 24TB from the survivors; the rating is a ceiling and not a measured rate, but it explains why rebuilds fail late and why RAID 6 is the default for big drives. What makes the job recoverable is what the rebuild left behind: the drive that failed first still holds every stripe the rebuild never reached, and the survivor that stopped it usually reads on a bench with its bad areas last. A set of two to four members is £500 + VAT upwards after the free look, fixed in writing, 5–10 days at the bench.

Free first lookOne fixed figure in writingNo data, no bill on most jobsReturn postage paid

Rather talk it through? An engineer answers the bench line
0800 6890668

Power down. Label every slot before a drive comes out. Do not rebuild, do not force a drive online, do not import or clear a foreign configuration, do not initialise, and do not run chkdsk, fsck, zpool import -F or mdadm --create on the members. A rebuild reads every sector of every survivor and writes to the replacement; each of the commands writes to the drives. On images, all of them are reversible. On the originals, none of them is.

What a rebuild reads, why it fails on large sets, and where the data still is.

A degraded RAID 5 of four drives has three survivors. A rebuild onto the replacement reads all three from end to end, computes the missing member's block for each stripe by XOR, and writes it. Every sector of every survivor is asked for, including the ones nobody has read since the array was built and the ones a survivor has been quietly reallocating around. A survivor with pending sectors, SMART attribute 197, has sectors it cannot currently give, and the rebuild asks for all of them. When it meets one, the controller either drops that survivor, which fails the virtual disk, or on a PERC skips the stripe and punctures it, finishing a rebuild with holes.

The unrecoverable-read-error rating is why large sets are exposed. Consumer drives are rated at one URE per 10^14 bits read, enterprise and nearline drives at one per 10^15; 10^14 bits is about 12.5TB. A rebuild on four 8TB drives reads about 24TB from the survivors, which on the naive reading of the rating is more than one expected error. The rating is a specification ceiling rather than a measured rate, and if it were a real rate most drives could not complete a long sequential read; a 2013 paper on reliability models put UREs between 10^-14 and 10^-15 per bit and computed that an eight-plus-two RAID 6 of 1TB drives had only about a 53 per cent chance of reading its survivors clean after two failures. The honest use of the number is the one the calculator makes: a reason rebuilds fail late, and a reason for RAID 6 on big drives, not a prediction for any one array.

On the bench the arithmetic is the same and the writes are not. The survivor that stopped the rebuild is imaged with its weak areas last, on hardware that retries the way a controller will not; the drive that failed first is imaged too, because it still holds every stripe the rebuild had not reached when it stopped, and on a set that was degraded for a week those are most of them. The set is assembled in software from the current members, with the first-dropped drive's image used only where a survivor genuinely cannot read, and the file system repaired on a copy of the virtual volume.

What it says, and what it means.

Describe yours to us →
What you see The usual reason Where that leaves you
Rebuild stalled at n per centA survivor's unrecoverable read errorPower down; image the survivor slowly; the first-dropped drive fills the rest
Rebuild finished with errors / puncturedStripes skippedImage before recreating; the first-dropped drive fills the holes
Second drive Failed; virtual disk FailedThe controller dropped the survivorEvery member imaged; assembled from images
Hot spare rebuilding onto a set with a predictive survivorAbout to failPull the power; image first
Rebuild ran from the stale driveWrong direction on a mirror or a forced-online memberStop; the current data may survive on the source
Rebuild restarted repeatedlyEach attempt re-reads the weak survivorStop restarting; each pass costs sectors

From the parcel arriving to your files going back.

Work we have closed →
01

Logged the day it lands, and the first look costs nothing Free

A case number goes on the parcel and a number on every member the day it is opened, matched to the slot you wrote on it. Each member goes on the imager its interface needs, never on a controller, and its metadata is read before a sector is: the level, the order, the stripe size, the event counts that say which member is current and which dropped out first. An engineer settles what has happened to the set and how much of it can honestly be read back. Back to you come two things together: a straight note of what is liftable and what is not, plus one figure, fixed and written down. Accept it, or decline and owe us nothing.

Nothing to pay for lookingA single figure, put in writingNo rebuilds, no imports, no initialise
02

Every member imaged, including the one that failed first

Every drive in the set is imaged sector by sector, weak areas last, on hardware that controls every retry, with a map of what could not be read kept for each. Members with failed heads go to the clean bench first; that is the drive site's work and the same bench does it. The drive that dropped out first is imaged too, because it still holds every stripe written before it dropped, and a rebuild that stalled part-way never reached them.

Sector by sector, weak areas lastThe first-dropped member included
03

The geometry, from the metadata or from the parity

A set that failed during a rebuild is assembled from three kinds of image: the survivors that are still current, the survivor that stopped the rebuild imaged with its weak areas last, and the first-dropped member. Parity across the images says which stripes the rebuild had recomputed onto the replacement, which the stopped survivor cannot give, and which the first-dropped drive still holds unchanged; each stripe is taken from the member that has it right.

Metadata first, parity secondOrder, stripe, rotation, offset
04

Assembled in software, and repaired on the virtual volume

The set is put together from the images in software, with nothing written to any of them: the current members in, the stale member used only to fill holes a survivor could not give. The file system is checked and repaired on a copy of the virtual volume, VMFS and CSV volumes opened and the virtual machines' disks extracted, databases repaired where they need it. The originals are not touched again.

Nothing written to the imagesVirtual machines and databases opened
05

You see the file list before you pay

What was recovered is listed for you first, and only then does a bill exist. Approve the list and it is invoiced; turn it down and it is not — and where nothing has come back, most jobs carry no charge at all. Recovered data travels home on fresh media bought in for your job, with the postage at our end. Your case is not closed until you have opened the files on a machine of your own.

No charge until you accept the figureFresh media, supplied with the job5–10 days at the bench

From the bench

  • Keep the first-dropped drive. It is not rubbish; it holds the stripes the rebuild never reached. The warranty courier must not take it.
  • Do not restart the rebuild. Each attempt re-reads the weak survivor from the start, and each read costs it sectors.
  • Send the part-written replacement too. The stripes the rebuild did complete are on it, correctly.
  • The calculator says why, not what will happen to yours. Use it to decide RAID 6 next time.

Tonight, in this order: stop writes and do not reboot repeatedly; press nothing the controller offers (F2, Import, Clear, Force online, Rebuild); photograph the screen and label the slots; export the log without changing state (PERC TTYLOG or storcli show all, an SSA diagnostic report, mdadm --examine on every member, zpool import with no flags, Get-VirtualDisk); power down and send every member, plus the controller if it was replaced or holds a key. The first-response page has the reasons.

One job, followed all the way through.

UK · RAR-2026-0619JOB LOGGED ✓

A four-drive RAID 5 of 8TB IronWolf drives in a small server, one member failed, a replacement fitted, the rebuild stalled at 71 per cent, restarted twice, and the virtual disk Failed

All five drives came in, including the part-written replacement and the first-dropped member. The survivor that had stopped the rebuild had 300 pending sectors and was imaged with them last, most yielding on retry; the first-dropped drive was imaged directly. The set was assembled from the two clean survivors, the slow survivor's image and the first-dropped drive's image for the stripes past 71 per cent that the rebuild had never reached, and the ext4 volume mounted read-only. Everything but two files under the unreadable sectors came back.

99.99% of the volume recovered7 days here, and back by post
Illustrative example — replace with a genuine case

What helps, and what harms.

Do this much first

  • Power down at the first stall
  • Keep and send the first-dropped drive and the part-written replacement
  • Send every member, labelled by slot
  • Tell us the percentage it reached and how many times it was restarted

What sets us back

  • Restarting the rebuild
  • Letting the engineer take the failed drive
  • Recreating the virtual disk after a rebuild with errors
  • Rebuilding onto a set with a predictive survivor in the first place

Questions answered before you commit.

My rebuild failed at 60 per cent. Is the data gone?

Usually not. The survivor that stopped it is imaged with its bad areas last, and the drive that failed first still holds the stripes past 60 per cent that the rebuild never reached. The set is assembled from the images.

Why do rebuilds fail on large drives?

Because they read every sector of every survivor, and the rated unrecoverable-read-error figures make one error over 24TB or more of reads likely on paper. The rating is a ceiling and not a measured rate, but it is why RAID 6 is the default for big drives.

Should I restart the rebuild?

No. Each attempt re-reads the weak survivor from the start and costs it sectors. Power down and image it.

What does it cost?

£500 + VAT upwards for a set of two to four members after the free look, fixed in writing; larger sets from £1,250 + VAT.

How long does it take?

5–10 days at the bench.

Nothing gets worse while it is powered down.

Looking at it is free. Back comes a list of what opened and what did not, together with a single price to finish, set down in writing while you are still free to say no. Until that list reaches you, leave the server off and the drives in their slots.

0800 6890668