Server Diagnostics: Difference between revisions
(Fix Services section (function probe), rename SSD Wear, update Limitations) |
(docs(diagnostics): rewrite Software Versions as a readable field table) |
||
| (4 intermediate revisions by the same user not shown) | |||
| Line 69: | Line 69: | ||
|} | |} | ||
The overall verdict is the highest severity present: any '''critical''' → CRITICAL; else any '''warning''' → DEGRADED; else HEALTHY. The health index (0–100) shown in the report is computed as 100 minus 25 per critical, 8 per warning, and 1 per optimization, clamped to 0–100. A 100 means nothing was flagged. | The overall verdict is the highest severity present: any '''critical''' → CRITICAL; else any '''warning''' → DEGRADED; else HEALTHY. The health index (0–100) shown in the report is computed as 100 minus 25 per critical, 8 per warning, and 1 per optimization, clamped to 0–100. A 100 means nothing was flagged. Optimization findings come in two kinds: tuning notes, which appear in the findings table, and '''Upgrade Opportunities''' (category Upgrade), which the Portal renders in a separate banner with an HPCMATE service offer and a [[contact]]-sales call-to-action — see Upgrade Opportunities below. | ||
== Check Reference == | == Check Reference == | ||
| Line 124: | Line 124: | ||
* False positives: on consumer/prosumer cards that do not [[support]] ECC (for example many TITAN/GeForce-class GPUs), the counters read back as unavailable/zero — that is '''not a failure, it is simply not applicable''', and the check should be read as N/A on such cards. | * False positives: on consumer/prosumer cards that do not [[support]] ECC (for example many TITAN/GeForce-class GPUs), the counters read back as unavailable/zero — that is '''not a failure, it is simply not applicable''', and the check should be read as N/A on such cards. | ||
=== GPU MIG === | |||
* Trigger: the per-GPU MIG mode ('''nvidia-smi --query-gpu=mig.mode.current''') plus the GPU-instance table ('''nvidia-smi mig -lgi'''). | |||
* Why: MIG is the datacenter multi-tenant primitive — it slices one GPU into isolated instances with dedicated memory and compute. Knowing whether a card is sliced tells you its intended tenancy model. | |||
* Severity: info for MIG mode off (the default, whole-card use), MIG on with instances created, or MIG on but unsliced. '''Warning''' for a MIG-capable card that reports no MIG state at all (display mode). | |||
* Display mode: a workstation Blackwell card (for example RTX PRO 4000/5000/6000) in '''display mode''' reports mig.mode of [N/A] — byte-for-byte identical to a genuinely non-MIG card. The check distinguishes the two by GPU model name: a known MIG-capable model reading all-N/A is flagged as a warning, because MIG profiles are not exposed until the card is switched to compute mode (displaymodeselector --gpumode compute --auto, then reboot). | |||
* False positives: a truly non-MIG card (TITAN, GeForce) also reports [N/A], but it does not match the MIG-capable model list, so it produces '''no''' finding — the warning fires only when the model is known to support MIG but the mode is hidden behind display mode. | |||
=== GPU XID Errors === | === GPU XID Errors === | ||
| Line 203: | Line 210: | ||
=== Software Versions === | === Software Versions === | ||
A read-only snapshot of the software running on the box. It never changes anything — it only records versions so an operator can see what is installed and decide what to update or re-flash. | |||
* Why: | |||
* Severity: info ( | {| class="wikitable" | ||
* False positives: a | !! Item | ||
!! What it reports | |||
!! Typical value | |||
|- | |||
| OS | |||
| Distribution and release, as the OS identifies itself | |||
| Rocky [[Linux]] 9.4 | |||
|- | |||
| Kernel | |||
| The running kernel version | |||
| 5.14.0-427.el9.x86_64 | |||
|- | |||
| Uptime | |||
| How long the host has been up | |||
| up 3 days, 4 hours | |||
|- | |||
| GPU driver | |||
| Installed [[NVIDIA driver]] version (first GPU) | |||
| 550.54.15 | |||
|- | |||
| NVMe firmware | |||
| Firmware revision of each NVMe drive | |||
| 0x114 | |||
|- | |||
| Docker | |||
| Container engine version | |||
| 26.1.4 | |||
|- | |||
| Python | |||
| Python 3 interpreter version | |||
| 3.11.2 | |||
|- | |||
| Node.js | |||
| Node.js runtime version | |||
| v22.0.0 | |||
|- | |||
| CUDA | |||
| CUDA toolkit release | |||
| 12.2 | |||
|- | |||
| GCC | |||
| Compiler version | |||
| 12.3.0 | |||
|- | |||
| smartctl | |||
| Smartmontools version (disk health tooling) | |||
| 7.3 | |||
|} | |||
Notes: | |||
* Severity: info — a reference snapshot, not a fault. It never changes the health verdict. | |||
* Conditional fields: GPU driver, NVMe firmware, and the package rows appear only when that software is actually present on the host; a minimal box shows just OS, kernel, and uptime. | |||
* In the Portal: the finding shows the versions as a compact key=value list (first eight); the full list is under the 'Raw diagnostic data' expander. | |||
* False positives: an old-but-stable driver or firmware is not broken — cross-reference the recorded versions against vendor advisories before acting. This is a planning aid, not a pass/fail check. | |||
=== Hardware Identity === | |||
* Trigger: system / baseboard / [[BIOS]] identity from '''dmidecode''' (with a '''/sys/class/dmi/id''' fallback when root is unavailable), BMC info from '''sudo ipmitool mc info''' ([[IPMI]]), full FRU inventory from '''sudo ipmitool fru''' (Board Mfg/Serial, Product Mfg/Serial, chassis), and every NIC MAC address (loopback excluded). | |||
* Why: the board model, [[BIOS]] revision, BMC firmware, FRU serials, and NIC MACs are the asset identity used for tracking, [[warranty]]/RMA matching, and firmware upgrade planning. Recording them in the report turns the box into a known asset with a concrete upgrade target. | |||
* Severity: info (a reference snapshot, never a fault). | |||
* Summary line: the finding detail is a single pipe-joined identity, for example '''Supermicro H12SSL-NT v1.01 | BIOS 2.3 | BMC 1.00 | SN 0123456789 | 3 NICs''' — vendor + board model (+ board version), [[BIOS]] version, BMC firmware, the first available serial (DMID system serial, falling back to FRU product/board serial), and the count of '''physical''' NICs. | |||
* Physical NIC count: virtual interfaces (docker bridge, veth pairs, tap/tun, kube/cni/calico/flannel, wireguard) are excluded from the NIC count so a Docker host does not read as "17 NICs" — only real onboard/attached NICs are counted. | |||
* Privilege: the BMC and FRU reads require root (local-bus [[IPMI]] is root-only), so the check runs them with '''sudo'''. On a host where the SSH user has no passwordless sudo (factory provisioning missing), those fields are simply absent — the check degrades to dmidecode/sysfs identity rather than faulting. | |||
* False positives: none by design. If '''dmidecode''' cannot be run, the check falls back to '''/sys/class/dmi/id''' fields; if there is no BMC (consumer/workstation board) the BMC/FRU fields are empty and skipped. | |||
=== Power Health === | |||
* Trigger: [[IPMI]] '''chassis status''' (power overload / main power fault / cooling-fan fault / drive fault), voltage sensors from the SDR, PSU sensors (used to detect PMBus availability), and power-related events from the SEL. | |||
* Why: a power-subsystem fault (overload, main power fault, a failing cooling loop) is an immediate physical risk, and the SEL is the historical record of power events — together they name whether the power delivery is healthy. | |||
* Severity: '''critical''' for any chassis power fault; '''warning''' for an unexpected system power state or for power-related SEL events; '''info''' for voltage / PSU telemetry. | |||
* PMBus limitation: many boards (for example the MSI D3060GB2N) do not wire a PMBus/IPMB PSU telemetry bus, so the PSU sensors read "No Reading". The check reports this as '''info, not a fault''' — power is still verified through the voltage sensors and the chassis status. For long-term PSU wear monitoring, use a board whose PSU exposes PMBus. | |||
* False positives: "0 PSU No Reading (no PMBus)" is a board limitation and is never escalated to a fault. The critical path requires the BMC to explicitly report a power fault flag set to true. | |||
=== Upgrade Opportunities === | |||
* Trigger: a separate layer of checks that fires when a component has headroom a customer could buy more of — not when something is broken. Eight checks: storage (any real mount, LVM volume, or ZFS pool at 80% or more), RAM (80% or more in use), GPU VRAM (80% or more on any card) or MIG disabled on a MIG-capable datacenter GPU, NIC (fastest physical link at 10 GbE or below), CPU (16 cores or fewer, or a previous-generation part), cooling (sustained GPU temperature 75 degrees C or above), PSU (a single unit with no N+1 redundancy), and software (a bare OS with no GPU driver, [[CUDA]], or Python environment). | |||
* Why: these are revenue-relevant. The Portal surfaces them to the customer as an "Upgrade Opportunities" banner — a card per item, each carrying an HPCMATE service offer and a contact-sales call-to-action — so the same diagnostic that reassures a healthy machine also points the customer at the next thing worth buying. | |||
* Severity: all eight are '''optimization''' (weight 1) and are tagged '''category Upgrade'''. They never raise the overall verdict and are excluded from the fault table; they render only in their own banner above the findings. | |||
* False positives: the GPU MIG check is gated on a model-name allowlist (A100/A30/A40, H100/H200/H800, L40/L40S, B100/B200 and similar) because consumer cards (for example the TITAN RTX) have no MIG capability and must never be offered an upgrade they cannot take. The CPU check uses a conservative generation heuristic, so a well-chosen modern low-core part is the main case it can over-flag. A 1 GbE NIC on a deliberately single-tenant inference box is a legitimate configuration even though the check still surfaces the 25/100 GbE option. | |||
== Limitations and Common False Positives == | == Limitations and Common False Positives == | ||
Latest revision as of 13:29, 28 September 2026
Server Diagnostics
Overview
The HPCMATE Portal runs a read-only system diagnostic over SSH to produce a structured health report for any registered server. It never modifies the target — it connects, runs a battery of non-destructive checks, and scores the result.
Summary
- What: a deep, read-only health check that connects over SSH and inspects CPU, memory, disks, GPU, network, services, and the kernel log.
- Why: it turns raw system state into a single verdict (HEALTHY / DEGRADED / CRITICAL) plus per-item findings and recommendations, so a technician can see what is wrong and why it matters at a glance.
- When: before shipping a server, during pre-ship validation, or on demand from the Portal → Diagnostic view.
Purpose
- Goal: give operators and customers a defensible, explainable health assessment.
- Scope: hardware health (GPU, disks, memory, thermals, network) and software/service state on a single host, over SSH.
- Non-goals: it does not install, repair, reboot, or change configuration. Multi-node cluster health is out of scope.
How It Works
The Portal spawns a Python diagnostic over the serversetup venv (paramiko). It opens an SSH session to the target and runs a fixed set of shell probes. Every probe that crosses a threshold becomes a finding with a severity. Findings are aggregated into one overall verdict and a numeric health index. The whole run is read-only and typically takes a few seconds to a minute.
| Stage | What happens | Output |
|---|---|---|
| Connect | TCP reachability probe, then SSH authentication | reachable / unreachable |
| Collect | run the battery of read-only checks | raw measurements |
| Evaluate | compare each measurement to thresholds | findings (severity + detail + recommendation) |
| Score | roll findings up into one verdict + health index | HEALTHY / DEGRADED / CRITICAL |
Verdict Model
Every finding carries one of four severities:
| Severity | Meaning | Weight |
|---|---|---|
| critical | Hardware failure or data-integrity risk; act now | 25 |
| warning | Degrading or out-of-policy; schedule maintenance | 8 |
| optimization | Works fine, but a known improvement is available | 1 |
| info | Informational; no action required | 0 |
The overall verdict is the highest severity present: any critical → CRITICAL; else any warning → DEGRADED; else HEALTHY. The health index (0–100) shown in the report is computed as 100 minus 25 per critical, 8 per warning, and 1 per optimization, clamped to 0–100. A 100 means nothing was flagged. Optimization findings come in two kinds: tuning notes, which appear in the findings table, and Upgrade Opportunities (category Upgrade), which the Portal renders in a separate banner with an HPCMATE service offer and a contact-sales call-to-action — see Upgrade Opportunities below.
Check Reference
This is the explanation of why the diagnostic judges each item the way it does. Each check below names its trigger condition, why it matters, the severity it can raise, and the common false positives. The Portal report links each finding here.
CPU Load
- Trigger: the 15-minute average load is compared against the core count.
- Why: sustained load above the core count means work is queued for longer than it can be serviced; above 80% of capacity it starts to hurt interactivity.
- Severity: critical if the 15-min load exceeds the core count; warning if it is above 80%; otherwise info.
- False positives: a burst from a parallel job (MPI/NCCL) can briefly push the 1-minute load high even when the 15-minute load is nominal — the check deliberately uses the 15-minute value to avoid reacting to short bursts.
Memory
- Trigger: available RAM (from free -m) is compared to total.
- Why: low free memory means the kernel is close to OOM-killing workloads; high swap usage means RAM pressure is spilling to disk.
- Severity: critical if available is under 1 GB or under 10% of total; warning if under 20%; otherwise info. Swap over 50% used is a warning.
- False positives: on a machine running GPU inference, the page cache and buffer cache look like "used" memory but are reclaimable — the check uses the available column, not used, precisely so cached memory is not misread as pressure.
Swap
- Trigger: swap used over 50% of swap total.
- Why: sustained swapping degrades performance and is usually a sign the host is undersized for the workload.
- Severity: warning.
- False positives: a host with a very large swap partition may use a little swap for cold pages without any real pressure — watch the trend, not a single sample.
Disk Space
- Trigger: usage percentage of the filesystems mounted at /, /home, /var, /opt, /data.
- Why: a near-full filesystem can stop services, fail job writes, or wedge logs.
- Severity: critical at 95% or more; warning at 85% or more.
- False positives: a tmpfs (tmp/ram) that is mostly in use is often intentional (scratch space) and is not flagged; only the listed real filesystems are scored.
Disk SMART
- Trigger: smartctl -H per physical disk, plus the Reallocated/Pending/Offline-Uncorrectable attribute counters.
- Why: SMART is the early-warning system for disk failure; uncorrectable errors mean data integrity is already at risk.
- Severity: critical if health is FAILED or there are any Offline_Uncorrectable sectors; warning for any Reallocated or Pending sectors; info on PASSED.
- False positives: smartctl missing or unable to read the disk reports UNKNOWN, not a failure — on hosts where smartmontools is not installed or the disk does not expose SMART, the check is simply not informative. A non-zero pending count can also be a transient that the disk corrects; monitor the trend before replacing.
GPU Temperature
- Trigger: per-GPU temperature from nvidia-smi.
- Why: sustained high temperature causes throttling and, over time, physical damage; it is a direct signal of inadequate cooling.
- Severity: critical above 90°C; warning above 80°C.
- False positives: a short heavy burst can push a card briefly into the warning band and settle back down; a single sample is a snapshot, so confirm it is sustained before treating it as a cooling fault. On a dev box running inference, an 80–90°C card under load is often expected, not a defect.
GPU ECC Errors
- Trigger: the lifetime corrected and uncorrectable ECC error counters from nvidia-smi.
- Why: uncorrectable ECC errors are a physical memory-cell failure — data the GPU computed may be wrong; a fast-growing corrected count is early degradation.
- Severity: critical for any uncorrectable error; warning above 1000 lifetime corrected.
- False positives: on consumer/prosumer cards that do not support ECC (for example many TITAN/GeForce-class GPUs), the counters read back as unavailable/zero — that is not a failure, it is simply not applicable, and the check should be read as N/A on such cards.
GPU MIG
- Trigger: the per-GPU MIG mode (nvidia-smi --query-gpu=mig.mode.current) plus the GPU-instance table (nvidia-smi mig -lgi).
- Why: MIG is the datacenter multi-tenant primitive — it slices one GPU into isolated instances with dedicated memory and compute. Knowing whether a card is sliced tells you its intended tenancy model.
- Severity: info for MIG mode off (the default, whole-card use), MIG on with instances created, or MIG on but unsliced. Warning for a MIG-capable card that reports no MIG state at all (display mode).
- Display mode: a workstation Blackwell card (for example RTX PRO 4000/5000/6000) in display mode reports mig.mode of [N/A] — byte-for-byte identical to a genuinely non-MIG card. The check distinguishes the two by GPU model name: a known MIG-capable model reading all-N/A is flagged as a warning, because MIG profiles are not exposed until the card is switched to compute mode (displaymodeselector --gpumode compute --auto, then reboot).
- False positives: a truly non-MIG card (TITAN, GeForce) also reports [N/A], but it does not match the MIG-capable model list, so it produces no finding — the warning fires only when the model is known to support MIG but the mode is hidden behind display mode.
GPU XID Errors
- Trigger: NVRM "Xid" lines from the kernel log (dmesg).
- Why: XID is NVIDIA's hardware fault code. Certain codes (48, 63, 64, 79) mean the GPU fell off the bus or hit a fatal fault; others (31, 43, 45, 68) are page-fault/ECC related.
- Severity: critical for the fatal set; warning for the page-fault set; otherwise info.
- False positives: a Xid can be logged for a transient context error that does not recur. The check reports the recent set so an operator can judge whether it is recurring; a single benign Xid does not by itself mean the card must be replaced.
Network Links
- Trigger: the state of each network interface from ip -brief link, excluding loopback and virtual (docker/veth/bridge) interfaces.
- Why: a physical NIC that is down means that path has no connectivity — a real problem for a server that is expected to be reachable on it.
- Severity: warning for each non-lo, non-virtual interface in the DOWN state.
- False positives: this is the check most often "wrong" for good reason. (a) A NIC that is intentionally not in use (a spare or a management port that is switched off) is legitimately down. (b) A USB NIC that is not plugged in shows as DOWN. (c) Some hosts run networking via a different stack (for example NetworkManager or a custom unit) where a legacy interface name is down by design. A DOWN flag here is a prompt to confirm whether that interface is supposed to be up, not a confirmed fault.
NIC Speed
- Trigger: negotiated link speed from ethtool per interface.
- Why: a link that auto-negotiated below the port's rating (for example 1 GbE on a 10 GbE port, or a bad cable/SFP) silently caps throughput.
- Severity: warning if a real interface negotiates below 1000 Mbps.
- False positives: virtual/USB interfaces report "Unknown" or auto and are skipped; a 10 GbE link that negotiates 1 GbE may be an intentional patch for a test — confirm the intent.
Temperature Sensors
- Trigger: chassis/CPU temperature readings from sensors (lm-sensors).
- Why: a sensor reading in the dangerous band means the cooling loop (fans, pump, airflow) is failing to hold the package temperature.
- Severity: critical above 95°C; warning above 85°C.
- False positives: many bare-metal servers do not expose a usable sensor map to lm-sensors, in which case the check reports nothing rather than a fault. A high reading on a heavily loaded node is expected and should be read together with the GPU/CPU load.
Services
- Trigger: for each expected service (sshd, docker, NetworkManager/network), the diagnostic first checks the systemd unit state (systemctl is-active). If the unit is active the service is considered working. If the unit is not active, a function probe determines whether the service is actually serving: sshd → is port 22 listening?; docker → does docker info succeed?; network → is there a default route and a usable interface?
- Why: a service that is genuinely down is a common root cause of "the server does not respond the way it should". However, systemd unit state alone is not a reliable health signal.
- Severity: warning only when the unit is not active AND the function probe confirms the service is not actually serving. If the unit reads inactive but the function probe succeeds (port listening, default route present, docker responding), the service is considered working and no finding is raised.
- False positives: this is the check most often "wrong" for good reason. (a) On Ubuntu 24.04, sshd and NetworkManager are commonly socket-activated — the long-running service unit reads as inactive while the socket accepts connections and the service is fully functional. (b) On dev/test hosts, sshd may run as a non-systemd daemon (a service script or a container) — both sshd.service and sshd.socket read inactive while port 22 is serving and you are literally connected over SSH. (c) The legacy network unit is inactive on hosts managed by NetworkManager. In all these cases the function probe confirms the service is working, so no warning is raised. A finding is only produced when the service is genuinely down (unit inactive and function probe fails).
Kernel Errors
- Trigger: the most recent dmesg messages at error/critical/alert level.
- Why: the kernel log is the last word on hardware faults — machine-check exceptions (CPU/RAM), disk I/O errors, PCIe AER errors, and OOM kills all surface here.
- Severity: critical for machine-check and disk I/O errors; warning for PCIe AER and OOM events; otherwise info.
- False positives: a few stale error lines from a past event can linger in the ring buffer. The check reads the most recent window so old, already-resolved events do not keep flagging a healthy host; if a recurring pattern appears, treat it as real.
NUMA
- Trigger: the socket/node count from numactl --hardware.
- Why: multi-socket hosts get a large latency penalty when a process is scheduled on a socket different from the memory it touches; being NUMA-aware removes it.
- Severity: optimization (informational) when more than one NUMA node is present.
- False positives: none by design — a single-socket host simply produces no finding. On a multi-socket HPC node this is a genuine tuning opportunity, not a defect.
SSD Wear
- Trigger: for NVMe, the SMART/Health Log via sudo smartctl -a (Percentage Used, Data Units Written, Available Spare, Media/Data Integrity Errors, temperature); for SATA SSDs, the wear attributes.
- Why: every SSD has a finite write budget (TBW/dWPD). "Percentage Used" is the vendor's normalized wear gauge; crossing it means the remaining life can no longer guarantee the rated endurance. Data integrity errors are a direct signal of media failure.
- Severity: critical at 80% or more wear, or any media/integrity error; warning at 60% or more; otherwise info (reports wear, total written, spare capacity, and temperature).
- False positives: on a lightly used server a 2% gauge with low total-bytes-written is normal and healthy. The gauge is not "health" — a drive at 95% wear can still pass a short SMART test, which is exactly why the write-based gauge matters more than the pass/fail flag.
Network Health
- Trigger: per-interface error and drop counters from /proc/net/dev (receive/ transmit errors and drops), excluding loopback and virtual interfaces.
- Why: physical-layer problems (bad cable, dirty/failing SFP, a dying NIC) show up as a growing error/drop counter long before the link goes down. Catching this is what turns a "mystery slowness" into a replace-this-part answer.
- Severity: critical for sustained high error counts (thousands or more); warning for any non-zero errors or a high drop count; otherwise no finding.
- False positives: a small, stable drop count that does not grow is usually harmless (a brief boot-time burst). The counters are cumulative since boot, so the meaningful signal is whether they keep climbing — compare against a previous run before replacing hardware. Virtual interfaces (tap/tun/veth) are excluded because their counters carry different meanings.
Syslog Patterns
- Trigger: the last week of error-priority journal (journalctl --priority=err), categorized into patterns: machine-check, disk I/O, PCIe AER, OOM, GPU/NVIDIA, USB, NIC link, NVMe.
- Why: hardware faults repeat and leave a signature. A cluster of machine-check events points at CPU/RAM; I/O errors point at a specific disk; PCIe AER at a GPU or riser; OOM at memory pressure. Grouping the log tells you which component to replace, not just that something is wrong.
- Severity: critical for multiple machine-check or disk I/O events; warning for a recurring PCIe-AER, GPU, NIC, or OOM pattern; info only for high overall log volume with no dominant pattern.
- False positives: a large volume of unrelated error-level messages (for example repetitive protocol chatter) raises the "log volume" info note without being a fault. The check deliberately escalates only on a recognizable hardware pattern, not on raw message count.
Storage Space
- Trigger: df -h over all real filesystems (NFS included) plus an inode check (df -i).
- Why: a full filesystem stops writes and can take services down; a full inode table stops file creation even when space remains. Both are capacity-planning signals that lead directly to an expansion or a cleanup.
- Severity: critical at 95% capacity; warning at 85% or at 90% of inodes; info at 70% (early capacity planning).
- False positives: a near-full tmpfs or a deliberately small dedicated partition can trip the capacity thresholds without being a real risk — the check scores the major real filesystems and reports the mount so the operator can judge intent.
Software Versions
A read-only snapshot of the software running on the box. It never changes anything — it only records versions so an operator can see what is installed and decide what to update or re-flash.
| ! Item | ! What it reports | ! Typical value |
|---|---|---|
| OS | Distribution and release, as the OS identifies itself | Rocky Linux 9.4 |
| Kernel | The running kernel version | 5.14.0-427.el9.x86_64 |
| Uptime | How long the host has been up | up 3 days, 4 hours |
| GPU driver | Installed NVIDIA driver version (first GPU) | 550.54.15 |
| NVMe firmware | Firmware revision of each NVMe drive | 0x114 |
| Docker | Container engine version | 26.1.4 |
| Python | Python 3 interpreter version | 3.11.2 |
| Node.js | Node.js runtime version | v22.0.0 |
| CUDA | CUDA toolkit release | 12.2 |
| GCC | Compiler version | 12.3.0 |
| smartctl | Smartmontools version (disk health tooling) | 7.3 |
Notes:
- Severity: info — a reference snapshot, not a fault. It never changes the health verdict.
- Conditional fields: GPU driver, NVMe firmware, and the package rows appear only when that software is actually present on the host; a minimal box shows just OS, kernel, and uptime.
- In the Portal: the finding shows the versions as a compact key=value list (first eight); the full list is under the 'Raw diagnostic data' expander.
- False positives: an old-but-stable driver or firmware is not broken — cross-reference the recorded versions against vendor advisories before acting. This is a planning aid, not a pass/fail check.
Hardware Identity
- Trigger: system / baseboard / BIOS identity from dmidecode (with a /sys/class/dmi/id fallback when root is unavailable), BMC info from sudo ipmitool mc info (IPMI), full FRU inventory from sudo ipmitool fru (Board Mfg/Serial, Product Mfg/Serial, chassis), and every NIC MAC address (loopback excluded).
- Why: the board model, BIOS revision, BMC firmware, FRU serials, and NIC MACs are the asset identity used for tracking, warranty/RMA matching, and firmware upgrade planning. Recording them in the report turns the box into a known asset with a concrete upgrade target.
- Severity: info (a reference snapshot, never a fault).
- Summary line: the finding detail is a single pipe-joined identity, for example Supermicro H12SSL-NT v1.01 | BIOS 2.3 | BMC 1.00 | SN 0123456789 | 3 NICs — vendor + board model (+ board version), BIOS version, BMC firmware, the first available serial (DMID system serial, falling back to FRU product/board serial), and the count of physical NICs.
- Physical NIC count: virtual interfaces (docker bridge, veth pairs, tap/tun, kube/cni/calico/flannel, wireguard) are excluded from the NIC count so a Docker host does not read as "17 NICs" — only real onboard/attached NICs are counted.
- Privilege: the BMC and FRU reads require root (local-bus IPMI is root-only), so the check runs them with sudo. On a host where the SSH user has no passwordless sudo (factory provisioning missing), those fields are simply absent — the check degrades to dmidecode/sysfs identity rather than faulting.
- False positives: none by design. If dmidecode cannot be run, the check falls back to /sys/class/dmi/id fields; if there is no BMC (consumer/workstation board) the BMC/FRU fields are empty and skipped.
Power Health
- Trigger: IPMI chassis status (power overload / main power fault / cooling-fan fault / drive fault), voltage sensors from the SDR, PSU sensors (used to detect PMBus availability), and power-related events from the SEL.
- Why: a power-subsystem fault (overload, main power fault, a failing cooling loop) is an immediate physical risk, and the SEL is the historical record of power events — together they name whether the power delivery is healthy.
- Severity: critical for any chassis power fault; warning for an unexpected system power state or for power-related SEL events; info for voltage / PSU telemetry.
- PMBus limitation: many boards (for example the MSI D3060GB2N) do not wire a PMBus/IPMB PSU telemetry bus, so the PSU sensors read "No Reading". The check reports this as info, not a fault — power is still verified through the voltage sensors and the chassis status. For long-term PSU wear monitoring, use a board whose PSU exposes PMBus.
- False positives: "0 PSU No Reading (no PMBus)" is a board limitation and is never escalated to a fault. The critical path requires the BMC to explicitly report a power fault flag set to true.
Upgrade Opportunities
- Trigger: a separate layer of checks that fires when a component has headroom a customer could buy more of — not when something is broken. Eight checks: storage (any real mount, LVM volume, or ZFS pool at 80% or more), RAM (80% or more in use), GPU VRAM (80% or more on any card) or MIG disabled on a MIG-capable datacenter GPU, NIC (fastest physical link at 10 GbE or below), CPU (16 cores or fewer, or a previous-generation part), cooling (sustained GPU temperature 75 degrees C or above), PSU (a single unit with no N+1 redundancy), and software (a bare OS with no GPU driver, CUDA, or Python environment).
- Why: these are revenue-relevant. The Portal surfaces them to the customer as an "Upgrade Opportunities" banner — a card per item, each carrying an HPCMATE service offer and a contact-sales call-to-action — so the same diagnostic that reassures a healthy machine also points the customer at the next thing worth buying.
- Severity: all eight are optimization (weight 1) and are tagged category Upgrade. They never raise the overall verdict and are excluded from the fault table; they render only in their own banner above the findings.
- False positives: the GPU MIG check is gated on a model-name allowlist (A100/A30/A40, H100/H200/H800, L40/L40S, B100/B200 and similar) because consumer cards (for example the TITAN RTX) have no MIG capability and must never be offered an upgrade they cannot take. The CPU check uses a conservative generation heuristic, so a well-chosen modern low-core part is the main case it can over-flag. A 1 GbE NIC on a deliberately single-tenant inference box is a legitimate configuration even though the check still surfaces the 25/100 GbE option.
Limitations and Common False Positives
The diagnostic is a snapshot, not a continuous monitor. Read a single run that way: a transient (a load burst, a one-off Xid, a momentarily hot GPU) can show up even on a healthy machine. The checks are tuned to reduce the noisiest false positives: memory uses the available column (not "used") so page cache is not misread as pressure; CPU load uses the 15-minute average (not 1-minute) so parallel-job bursts do not trigger; missing sensors or tools are reported as "not applicable" rather than "failed"; the Services check uses a function probe (port listening, default route, docker responding) so socket-activated or non-systemd services are not flagged as down; and the Services and Network checks together report a network condition only once. The remaining areas that most benefit from a human look are the GPU ECC check (N/A on cards without ECC support) and the SSD Wear gauge (a low gauge on a lightly used drive is healthy, not a fault). When a finding is not obvious, open the Why? link for that check and confirm with the suggested command before treating the host as faulty.
References
- HPCMATE Portal — Diagnostic view (POST/GET /api/servers/:id/diagnostic)
- diagnose_server.py — the read-only check implementation
- NVIDIA nvidia-smi documentation (temperature, ECC, XID)
- smartmontools (smartctl) documentation
- lm-sensors documentation