[ad_1]
I used HDSentinel for Unix (Ubuntu 24.10) to test an Intel SSD DC S4500 Series 1.92TB disk as follows:
root@stephen-All-Series:~# HDSentinel
Hard Disk Sentinel for LINUX console 0.20c-x64.10851 (c) 2024 [email protected]
Start with -r [reportfile] to save data to report, -h for help
Examining hard disk configuration ...
HDD Device 0: /dev/sda
HDD Model ID : INTEL SSDSC2KB019T7
HDD Serial No: PHYS737500DM1P9DGN
HDD Revision : SCV10121
HDD Size : 1831420 MB
Interface : S-ATA Gen3, 6 Gbps
Temperature : 21 °C
Highest Temp.: 21 °C
Health : 49 %
Performance : 100 %
Power on time: 1355 days, 22 hours
Est. lifetime: 112 days
Total written: 5,206.45 TB
The status of the solid state disk is PERFECT. Problematic or weak sectors were not found.
The health is determined by SSD specific S.M.A.R.T. attribute(s): #233 Media Wearout Indicator
It is recommended to backup often to prevent data loss.
The health was 49% as was attributed to the “#233 Media Wearout Indicator”.
Following up on that, I executed the following:
stephen@stephen-All-Series:~$ sudo smartctl -A /dev/sda
smartctl 7.4 2023-08-01 r5530 [x86_64-linux-6.11.0-9-generic] (local build)
Copyright (C) 2002-23, Bruce Allen, Christian Franke, www.smartmontools.org
=== START OF READ SMART DATA SECTION ===
SMART Attributes Data Structure revision number: 1
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
5 Reallocated_Sector_Ct 0x0032 100 100 000 Old_age Always - 0
9 Power_On_Hours 0x0032 100 100 000 Old_age Always - 32542
12 Power_Cycle_Count 0x0032 100 100 000 Old_age Always - 33
170 Available_Reservd_Space 0x0033 100 100 010 Pre-fail Always - 0
171 Program_Fail_Count 0x0032 100 100 000 Old_age Always - 0
172 Erase_Fail_Count 0x0032 100 100 000 Old_age Always - 0
174 Unsafe_Shutdown_Count 0x0032 100 100 000 Old_age Always - 31
175 Power_Loss_Cap_Test 0x0033 100 100 010 Pre-fail Always - 2709 (225 7757)
183 SATA_Downshift_Count 0x0032 100 100 000 Old_age Always - 0
184 End-to-End_Error_Count 0x0033 100 100 090 Pre-fail Always - 0
187 Uncorrectable_Error_Cnt 0x0032 100 100 000 Old_age Always - 0
190 Drive_Temperature 0x0022 079 079 000 Old_age Always - 21 (Min/Max 17/21)
192 Unsafe_Shutdown_Count 0x0032 100 100 000 Old_age Always - 31
194 Temperature_Celsius 0x0022 100 100 000 Old_age Always - 21
197 Pending_Sector_Count 0x0012 100 100 000 Old_age Always - 36
199 CRC_Error_Count 0x003e 100 100 000 Old_age Always - 0
225 Host_Writes_32MiB 0x0032 100 100 000 Old_age Always - 170604988
226 Workld_Media_Wear_Indic 0x0032 100 100 000 Old_age Always - 52961
227 Workld_Host_Reads_Perc 0x0032 100 100 000 Old_age Always - 59
228 Workload_Minutes 0x0032 100 100 000 Old_age Always - 1952266
232 Available_Reservd_Space 0x0033 100 100 010 Pre-fail Always - 0
233 Media_Wearout_Indicator 0x0032 049 049 000 Old_age Always - 0
234 Thermal_Throttle_Status 0x0032 100 100 000 Old_age Always - 0/0
241 Host_Writes_32MiB 0x0032 100 100 000 Old_age Always - 170604988
242 Host_Reads_32MiB 0x0032 100 100 000 Old_age Always - 246119347
243 NAND_Writes_32MiB 0x0032 100 100 000 Old_age Always - 186350039
Following up on that, I referred to the Wikipedia article Self-Monitoring, Analysis and Reporting Technology which wrote:
Accuracy
A field study at Google[9] covering over 100,000 consumer-grade drives
from December 2005 to August 2006 found correlations between certain
S.M.A.R.T. information and annualized failure rates:In the 60 days following the first uncorrectable error on a drive
(S.M.A.R.T. attribute 0xC6 or 198) detected as a result of an offline
scan, the drive was, on average, 39 times more likely to fail than a
similar drive for which no such error occurred. First errors in
reallocations, offline reallocations (S.M.A.R.T. attributes 0xC4 and
0x05 or 196 and 5) and probational counts (S.M.A.R.T. attribute 0xC5
or 197) were also strongly correlated to higher probabilities of
failure. Conversely, little correlation was found for increased
temperature and no correlation for usage level. However, the research
showed that a large proportion (56%) of the failed drives failed
without recording any count in the “four strong S.M.A.R.T. warnings”
identified as scan errors, reallocation count, offline reallocation,
and probational count. Further, 36% of failed drives did so without
recording any S.M.A.R.T. error at all, except the temperature, meaning
that S.M.A.R.T. data alone was of limited usefulness in anticipating
failures.[10]
And also:
233 0xE9 Media Wearout Indicator (SSDs) or Power-On Hours Intel SSDs
report a normalized value from 100, a new drive, to a minimum of 1. It
decreases while the NAND erase cycles increase from 0 to the
maximum-rated cycles. Previously (pre-2010) occasionally used for
Power-On Hours (more typically reported in 0x09).
This solid state disk is only four years old (from December 2024), so Power-On Hours are not part of this parameter, which reports 49, exactly corresponding to the rated health of the disk in Linux HDSentinel of 49%.
My guess is that this disk received heavy usage in a server configuration and was eventually swapped out to reduce the chance of the server configuration going down for a new disk. After all:
Total written: 5,206.45 TB
from the first quote. That is a very large amount of written data. But it seems then that this disk should be fine for consumer applications which do not write so much data going forward. This disk is very large so if small writes are applied over a very large area (like in the concept of not just RAID 0 striping, but also other levels of RAID that give not just the benefit of speed – and being spread out on the disk but also the benefit of built-in backup), it seems that there would be very little chance that this disk would actually fail in a consumer application. Especially with regular backup to another drive before each power down, it seems like a very solid configuration.
I have not seen RAID striping on a single disk yet for Ubuntu Solid State Disks.
Is it possible to force striping across the SSD disk (RAID or otherwise) so that there is minimal problem of overwriting the same SSD NAND memory area? (Is it automatic with this disk)?
Is this rating really a problem for consumers with minimal disk usage (megabytes of usage per month at most), or is this more just a difficulty for industrial customers with high-bandwidth industrial servers (resulting in the reported 5,206.45 TB already written)?
Relevant Links
Interesting: Is over-provisioning of an SSD possible with dual-boot?
The article Ubuntu OS on SSD and RAID-5 for data files seems promising, but still missing is how to accomplish this with one SSD with Ubuntu.
|
Sources 2/ https://askubuntu.com/questions/1535334/is-49-device-health-from-233-media-wearout-indicator-for-an-intel-ssd-dc-s54 The mention sources can contact us to remove/changing this article |
[ad_2]