r/vmware
VMware 8.02 + Micron 7450 Pro 960GB NVMe M.2 2280
- upvotes
- 4
- comments
- 9
Post
Greetings! I'm attempting, and failing, to debug a spontaneous host reboot. My DevOps/SRE team strongly believe the origin of this host reboot is from a bad drive. I'd like to prove them wrong, or be proven wrong myself, and hope I'm in the right place. Background The configuration is small as far as servers go, Asrock Rack B550D4ID-2L2T Ryzen 9 5950X (16-Core) 64 GB Memory On a UPS that is honestly way too large (can provide *hours* of uptime) Prior to installing the new drive, after roughly 380 days of uptime, the host spontaneously rebooted. Approximately a month after the first time, it rebooted again. Once was an abnormality, twice is a pattern. Reviewing the logs, and the hardware, suggested that it was a drive issue (22110). Reviewing the hardware identified two problems; the drive wasn't on the HCL the drive was a 22110 (the 2L2T doesn't natively support 22110 M2s) Okay, I'm convinced -- time to buy a new drive. I purchased a 7450 PRO off of ebay with minimal use (< 20 power cycles, < 1 TB writes/reads, essentially perfect SMART metrics). I verified as much, and immediately attempted to install the drive. Issue My experience with the 7450 Pro drive + ESXI 8.02 has not been pleasant. Here is what we needed to get it to boot/install correctly; update disk to 512b physical sector format (thanks to another Reddit Post) Post installation and VM migration...
Keep reading with a free account
The rest of this post, and every signal for Micron, is in your free account.
Extracted from these lines
My DevOps/SRE team is well convinced that the disk is bad at this point (controller channel, not data channel). These are the logs that convinced them.
From the post
WARNING: HPP: HppNvmeThrottleLogForDevice:599: NVMe Cmd 0x1 (0x4579050d3840, 1050659) to dev "t10.NVMe____Micron_7450_MTFDKBA960TFR_______________8C2E02480175A000" on path "vmhba0:C0:T0:L0" Failed:
From the post
[comment u/L0rd_OverKill] There is a bug that sounds just like this for those drives. When they reach 65535 hours there is an overflow error and they panic. There is a firmware update from Micron