SuperServer · CSE-826 · CSE-846 · AOC-S3108 · Cisco UCS C220 · C240 · 12G SAS Modular RAID · Primergy RX2540 · PRAID EP420i · EP540i
Supermicro, Cisco and Fujitsu servers array recovery. Three server makers, one controller family, and the HBAs that hand the drives to software.
Supermicro's SuperServers and CSE-826, 846 and 847 storage chassis, Cisco's UCS C220 and C240, and Fujitsu's Primergy RX and TX servers all run Broadcom MegaRAID silicon under their own names: Supermicro's AOC-S3108 and S3916, Cisco's 12G SAS Modular RAID, Fujitsu's PRAID EP420i and EP540i. The states are storcli's, the metadata is DDF near the end of the members, and the MegaRAID page's dialect applies. Just as often the same servers carry a SAS HBA in IT mode, Supermicro's AOC-S3008 or a 9300-8i, passing the drives straight to ZFS, mdadm or Storage Spaces, in which case the software pages apply. This page maps each server line to what is underneath, and the array is reassembled from images the same way either way. A set of two to four members is £500 + VAT upwards after the free look, fixed in writing; large chassis from £1,250 + VAT.
Rather talk it through? An engineer answers the bench line
0800 6890668
Models and versions we see.
Which level is it →| Family | Models or versions | Metadata, defaults and notes |
|---|---|---|
| Supermicro RAID | AOC-S3108, S3916 | MegaRAID; storcli |
| Supermicro HBA | AOC-S3008, 9300-8i in IT mode | ZFS, mdadm; the software pages |
| Cisco UCS | 12G SAS Modular RAID, UCSC-RAID-M5 | MegaRAID; UCS Manager reports |
| Fujitsu Primergy | PRAID EP420i, EP540i, CP400i | MegaRAID; ServerView reports |
What tends to go wrong on these servers.
Supermicro, Cisco and Fujitsu servers symptoms, and how long each gives you.
Not listed? Describe it on the form →What the message means on these servers.
Describe yours to us →| What you see | The usual reason | Where that leaves you |
|---|---|---|
| Array degraded on the host's controller | A member out | Image the weak drives before rebuilding |
| Array failed; datastore or volume gone | Too many members out | Every member imaged; the array, then the volume, then the guests |
| Datastore or CSV inaccessible; array fine | The volume layer | A file-system job on the virtual volume |
| Virtual machine will not start | A torn or missing virtual disk | Extracted from the volume and repaired |
From the parcel arriving to your files going back.
Work we have closed →Logged the day it lands, and the first look costs nothing Free
A case number goes on the parcel and a number on every member the day it is opened, matched to the slot you wrote on it. Each member goes on the imager its interface needs, never on a controller, and its metadata is read before a sector is: the level, the order, the stripe size, the event counts that say which member is current and which dropped out first. An engineer settles what has happened to the set and how much of it can honestly be read back. Back to you come two things together: a straight note of what is liftable and what is not, plus one figure, fixed and written down. Accept it, or decline and owe us nothing.
Every member imaged, including the one that failed first
Every drive in the set is imaged sector by sector, weak areas last, on hardware that controls every retry, with a map of what could not be read kept for each. Members with failed heads go to the clean bench first; that is the drive site's work and the same bench does it. The drive that dropped out first is imaged too, because it still holds every stripe written before it dropped, and a rebuild that stalled part-way never reached them.
The geometry, from the metadata or from the parity
Where the controller's metadata survives on the members, the order, stripe size, parity rotation and data offset are read from it. Where it was cleared or overwritten, they are recovered from the data: parity across the members at the same offset should XOR to zero on a consistent stripe, which confirms the level and finds the stale member; where parity lands says the rotation; entropy at stripe edges and the file system's own anchors give the stripe size, the order and the start.
Assembled in software, and repaired on the virtual volume
The set is put together from the images in software, with nothing written to any of them: the current members in, the stale member used only to fill holes a survivor could not give. The file system is checked and repaired on a copy of the virtual volume, VMFS and CSV volumes opened and the virtual machines' disks extracted, databases repaired where they need it. The originals are not touched again.
You see the file list before you pay
What was recovered is listed for you first, and only then does a bill exist. Approve the list and it is invoiced; turn it down and it is not — and where nothing has come back, most jobs carry no charge at all. Recovered data travels home on fresh media bought in for your job, with the postage at our end. Your case is not closed until you have opened the files on a machine of your own.
From the bench
- Four layers: the array, the volume, the virtual disk, the guest. Each is opened in turn on the images, and each is a place a rival's page is silent.
- Do not let the hypervisor repair anything until the array is imaged. Its repair reads the survivors under load.
- Say which virtual machines matter. Extraction begins there.
What helps, and what harms.
Do this much first
- Export the array log and the hypervisor's view of the storage
- Power down and label every member by host and slot
- Send every member of every affected array or disk group
- Tell us which virtual machines or databases matter most
What sets us back
- Rebuilding the array with a weak survivor
- Running the hypervisor's repair or migration on a degraded store
- Formatting or re-adding a datastore
- Restoring a backup over the only copy
Questions answered before you commit.
Is my Cisco 12G RAID a MegaRAID?
Yes, with Cisco firmware and UCS Manager reporting the same states. The MegaRAID page's dialect applies, and the DDF on the drives reads the same on the bench.
Do I send the controller or the unit?
The controller only if it was replaced and the set is now foreign, or if it holds an encryption key. A NAS or DAS unit, only if its bridge holds the layout; the pages say which.
What does it cost?
A set of two to four members is £500 + VAT upwards after the free look; larger sets from £1,250 + VAT, fixed in writing.
How long does it take?
5–10 days at the bench for a set of two to four; 10–15 days at the bench for larger sets.
Begin here if yours is doing the same thing.
Nothing gets worse while it is powered down.
Looking at it is free. Back comes a list of what opened and what did not, together with a single price to finish, set down in writing while you are still free to say no. On most jobs an invoice only follows the data. Until that list reaches you, leave the server off and the drives in their slots.