Server Diagnostics

From HPCWIKI
Revision as of 19:41, 20 September 2026 by Clara (talk | contribs) (Fix Services section (function probe), rename SSD Wear, update Limitations)
Jump to navigation Jump to search

Server Diagnostics

Overview

The HPCMATE Portal runs a read-only system diagnostic over SSH to produce a structured health report for any registered server. It never modifies the target — it connects, runs a battery of non-destructive checks, and scores the result.

Summary

  • What: a deep, read-only health check that connects over SSH and inspects CPU, memory, disks, GPU, network, services, and the kernel log.
  • Why: it turns raw system state into a single verdict (HEALTHY / DEGRADED / CRITICAL) plus per-item findings and recommendations, so a technician can see what is wrong and why it matters at a glance.
  • When: before shipping a server, during pre-ship validation, or on demand from the Portal → Diagnostic view.

Purpose

  • Goal: give operators and customers a defensible, explainable health assessment.
  • Scope: hardware health (GPU, disks, memory, thermals, network) and software/service state on a single host, over SSH.
  • Non-goals: it does not install, repair, reboot, or change configuration. Multi-node cluster health is out of scope.

How It Works

The Portal spawns a Python diagnostic over the serversetup venv (paramiko). It opens an SSH session to the target and runs a fixed set of shell probes. Every probe that crosses a threshold becomes a finding with a severity. Findings are aggregated into one overall verdict and a numeric health index. The whole run is read-only and typically takes a few seconds to a minute.

Stage What happens Output
Connect TCP reachability probe, then SSH authentication reachable / unreachable
Collect run the battery of read-only checks raw measurements
Evaluate compare each measurement to thresholds findings (severity + detail + recommendation)
Score roll findings up into one verdict + health index HEALTHY / DEGRADED / CRITICAL

Verdict Model

Every finding carries one of four severities:

Severity Meaning Weight
critical Hardware failure or data-integrity risk; act now 25
warning Degrading or out-of-policy; schedule maintenance 8
optimization Works fine, but a known improvement is available 1
info Informational; no action required 0

The overall verdict is the highest severity present: any critical → CRITICAL; else any warning → DEGRADED; else HEALTHY. The health index (0–100) shown in the report is computed as 100 minus 25 per critical, 8 per warning, and 1 per optimization, clamped to 0–100. A 100 means nothing was flagged.

Check Reference

This is the explanation of why the diagnostic judges each item the way it does. Each check below names its trigger condition, why it matters, the severity it can raise, and the common false positives. The Portal report links each finding here.

CPU Load

  • Trigger: the 15-minute average load is compared against the core count.
  • Why: sustained load above the core count means work is queued for longer than it can be serviced; above 80% of capacity it starts to hurt interactivity.
  • Severity: critical if the 15-min load exceeds the core count; warning if it is above 80%; otherwise info.
  • False positives: a burst from a parallel job (MPI/NCCL) can briefly push the 1-minute load high even when the 15-minute load is nominal — the check deliberately uses the 15-minute value to avoid reacting to short bursts.

Memory

  • Trigger: available RAM (from free -m) is compared to total.
  • Why: low free memory means the kernel is close to OOM-killing workloads; high swap usage means RAM pressure is spilling to disk.
  • Severity: critical if available is under 1 GB or under 10% of total; warning if under 20%; otherwise info. Swap over 50% used is a warning.
  • False positives: on a machine running GPU inference, the page cache and buffer cache look like "used" memory but are reclaimable — the check uses the available column, not used, precisely so cached memory is not misread as pressure.

Swap

  • Trigger: swap used over 50% of swap total.
  • Why: sustained swapping degrades performance and is usually a sign the host is undersized for the workload.
  • Severity: warning.
  • False positives: a host with a very large swap partition may use a little swap for cold pages without any real pressure — watch the trend, not a single sample.

Disk Space

  • Trigger: usage percentage of the filesystems mounted at /, /home, /var, /opt, /data.
  • Why: a near-full filesystem can stop services, fail job writes, or wedge logs.
  • Severity: critical at 95% or more; warning at 85% or more.
  • False positives: a tmpfs (tmp/ram) that is mostly in use is often intentional (scratch space) and is not flagged; only the listed real filesystems are scored.

Disk SMART

  • Trigger: smartctl -H per physical disk, plus the Reallocated/Pending/Offline-Uncorrectable attribute counters.
  • Why: SMART is the early-warning system for disk failure; uncorrectable errors mean data integrity is already at risk.
  • Severity: critical if health is FAILED or there are any Offline_Uncorrectable sectors; warning for any Reallocated or Pending sectors; info on PASSED.
  • False positives: smartctl missing or unable to read the disk reports UNKNOWN, not a failure — on hosts where smartmontools is not installed or the disk does not expose SMART, the check is simply not informative. A non-zero pending count can also be a transient that the disk corrects; monitor the trend before replacing.

GPU Temperature

  • Trigger: per-GPU temperature from nvidia-smi.
  • Why: sustained high temperature causes throttling and, over time, physical damage; it is a direct signal of inadequate cooling.
  • Severity: critical above 90°C; warning above 80°C.
  • False positives: a short heavy burst can push a card briefly into the warning band and settle back down; a single sample is a snapshot, so confirm it is sustained before treating it as a cooling fault. On a dev box running inference, an 80–90°C card under load is often expected, not a defect.

GPU ECC Errors

  • Trigger: the lifetime corrected and uncorrectable ECC error counters from nvidia-smi.
  • Why: uncorrectable ECC errors are a physical memory-cell failure — data the GPU computed may be wrong; a fast-growing corrected count is early degradation.
  • Severity: critical for any uncorrectable error; warning above 1000 lifetime corrected.
  • False positives: on consumer/prosumer cards that do not support ECC (for example many TITAN/GeForce-class GPUs), the counters read back as unavailable/zero — that is not a failure, it is simply not applicable, and the check should be read as N/A on such cards.

GPU XID Errors

  • Trigger: NVRM "Xid" lines from the kernel log (dmesg).
  • Why: XID is NVIDIA's hardware fault code. Certain codes (48, 63, 64, 79) mean the GPU fell off the bus or hit a fatal fault; others (31, 43, 45, 68) are page-fault/ECC related.
  • Severity: critical for the fatal set; warning for the page-fault set; otherwise info.
  • False positives: a Xid can be logged for a transient context error that does not recur. The check reports the recent set so an operator can judge whether it is recurring; a single benign Xid does not by itself mean the card must be replaced.

Network Links

  • Trigger: the state of each network interface from ip -brief link, excluding loopback and virtual (docker/veth/bridge) interfaces.
  • Why: a physical NIC that is down means that path has no connectivity — a real problem for a server that is expected to be reachable on it.
  • Severity: warning for each non-lo, non-virtual interface in the DOWN state.
  • False positives: this is the check most often "wrong" for good reason. (a) A NIC that is intentionally not in use (a spare or a management port that is switched off) is legitimately down. (b) A USB NIC that is not plugged in shows as DOWN. (c) Some hosts run networking via a different stack (for example NetworkManager or a custom unit) where a legacy interface name is down by design. A DOWN flag here is a prompt to confirm whether that interface is supposed to be up, not a confirmed fault.

NIC Speed

  • Trigger: negotiated link speed from ethtool per interface.
  • Why: a link that auto-negotiated below the port's rating (for example 1 GbE on a 10 GbE port, or a bad cable/SFP) silently caps throughput.
  • Severity: warning if a real interface negotiates below 1000 Mbps.
  • False positives: virtual/USB interfaces report "Unknown" or auto and are skipped; a 10 GbE link that negotiates 1 GbE may be an intentional patch for a test — confirm the intent.

Temperature Sensors

  • Trigger: chassis/CPU temperature readings from sensors (lm-sensors).
  • Why: a sensor reading in the dangerous band means the cooling loop (fans, pump, airflow) is failing to hold the package temperature.
  • Severity: critical above 95°C; warning above 85°C.
  • False positives: many bare-metal servers do not expose a usable sensor map to lm-sensors, in which case the check reports nothing rather than a fault. A high reading on a heavily loaded node is expected and should be read together with the GPU/CPU load.

Services

  • Trigger: for each expected service (sshd, docker, NetworkManager/network), the diagnostic first checks the systemd unit state (systemctl is-active). If the unit is active the service is considered working. If the unit is not active, a function probe determines whether the service is actually serving: sshd → is port 22 listening?; docker → does docker info succeed?; network → is there a default route and a usable interface?
  • Why: a service that is genuinely down is a common root cause of "the server does not respond the way it should". However, systemd unit state alone is not a reliable health signal.
  • Severity: warning only when the unit is not active AND the function probe confirms the service is not actually serving. If the unit reads inactive but the function probe succeeds (port listening, default route present, docker responding), the service is considered working and no finding is raised.
  • False positives: this is the check most often "wrong" for good reason. (a) On Ubuntu 24.04, sshd and NetworkManager are commonly socket-activated — the long-running service unit reads as inactive while the socket accepts connections and the service is fully functional. (b) On dev/test hosts, sshd may run as a non-systemd daemon (a service script or a container) — both sshd.service and sshd.socket read inactive while port 22 is serving and you are literally connected over SSH. (c) The legacy network unit is inactive on hosts managed by NetworkManager. In all these cases the function probe confirms the service is working, so no warning is raised. A finding is only produced when the service is genuinely down (unit inactive and function probe fails).

Kernel Errors

  • Trigger: the most recent dmesg messages at error/critical/alert level.
  • Why: the kernel log is the last word on hardware faults — machine-check exceptions (CPU/RAM), disk I/O errors, PCIe AER errors, and OOM kills all surface here.
  • Severity: critical for machine-check and disk I/O errors; warning for PCIe AER and OOM events; otherwise info.
  • False positives: a few stale error lines from a past event can linger in the ring buffer. The check reads the most recent window so old, already-resolved events do not keep flagging a healthy host; if a recurring pattern appears, treat it as real.

NUMA

  • Trigger: the socket/node count from numactl --hardware.
  • Why: multi-socket hosts get a large latency penalty when a process is scheduled on a socket different from the memory it touches; being NUMA-aware removes it.
  • Severity: optimization (informational) when more than one NUMA node is present.
  • False positives: none by design — a single-socket host simply produces no finding. On a multi-socket HPC node this is a genuine tuning opportunity, not a defect.

SSD Wear

  • Trigger: for NVMe, the SMART/Health Log via sudo smartctl -a (Percentage Used, Data Units Written, Available Spare, Media/Data Integrity Errors, temperature); for SATA SSDs, the wear attributes.
  • Why: every SSD has a finite write budget (TBW/dWPD). "Percentage Used" is the vendor's normalized wear gauge; crossing it means the remaining life can no longer guarantee the rated endurance. Data integrity errors are a direct signal of media failure.
  • Severity: critical at 80% or more wear, or any media/integrity error; warning at 60% or more; otherwise info (reports wear, total written, spare capacity, and temperature).
  • False positives: on a lightly used server a 2% gauge with low total-bytes-written is normal and healthy. The gauge is not "health" — a drive at 95% wear can still pass a short SMART test, which is exactly why the write-based gauge matters more than the pass/fail flag.

Network Health

  • Trigger: per-interface error and drop counters from /proc/net/dev (receive/ transmit errors and drops), excluding loopback and virtual interfaces.
  • Why: physical-layer problems (bad cable, dirty/failing SFP, a dying NIC) show up as a growing error/drop counter long before the link goes down. Catching this is what turns a "mystery slowness" into a replace-this-part answer.
  • Severity: critical for sustained high error counts (thousands or more); warning for any non-zero errors or a high drop count; otherwise no finding.
  • False positives: a small, stable drop count that does not grow is usually harmless (a brief boot-time burst). The counters are cumulative since boot, so the meaningful signal is whether they keep climbing — compare against a previous run before replacing hardware. Virtual interfaces (tap/tun/veth) are excluded because their counters carry different meanings.

Syslog Patterns

  • Trigger: the last week of error-priority journal (journalctl --priority=err), categorized into patterns: machine-check, disk I/O, PCIe AER, OOM, GPU/NVIDIA, USB, NIC link, NVMe.
  • Why: hardware faults repeat and leave a signature. A cluster of machine-check events points at CPU/RAM; I/O errors point at a specific disk; PCIe AER at a GPU or riser; OOM at memory pressure. Grouping the log tells you which component to replace, not just that something is wrong.
  • Severity: critical for multiple machine-check or disk I/O events; warning for a recurring PCIe-AER, GPU, NIC, or OOM pattern; info only for high overall log volume with no dominant pattern.
  • False positives: a large volume of unrelated error-level messages (for example repetitive protocol chatter) raises the "log volume" info note without being a fault. The check deliberately escalates only on a recognizable hardware pattern, not on raw message count.

Storage Space

  • Trigger: df -h over all real filesystems (NFS included) plus an inode check (df -i).
  • Why: a full filesystem stops writes and can take services down; a full inode table stops file creation even when space remains. Both are capacity-planning signals that lead directly to an expansion or a cleanup.
  • Severity: critical at 95% capacity; warning at 85% or at 90% of inodes; info at 70% (early capacity planning).
  • False positives: a near-full tmpfs or a deliberately small dedicated partition can trip the capacity thresholds without being a real risk — the check scores the major real filesystems and reports the mount so the operator can judge intent.

Software Versions

  • Trigger: a snapshot of the kernel, OS, GPU driver, NVMe firmware, and key packages (docker, python, nodejs, cuda, gcc, smartctl).
  • Why: an outdated GPU driver, an old NVMe firmware, or a missing critical security patch are maintenance items with known fixes. Recording the versions turns the report into a concrete upgrade/firmware-flash plan.
  • Severity: info (reference, not a fault).
  • False positives: a driver that is old but stable for the workload is not broken — this is a planning aid, not a pass/fail check. Cross-reference the recorded versions against vendor advisories before acting.

Limitations and Common False Positives

The diagnostic is a snapshot, not a continuous monitor. Read a single run that way: a transient (a load burst, a one-off Xid, a momentarily hot GPU) can show up even on a healthy machine. The checks are tuned to reduce the noisiest false positives: memory uses the available column (not "used") so page cache is not misread as pressure; CPU load uses the 15-minute average (not 1-minute) so parallel-job bursts do not trigger; missing sensors or tools are reported as "not applicable" rather than "failed"; the Services check uses a function probe (port listening, default route, docker responding) so socket-activated or non-systemd services are not flagged as down; and the Services and Network checks together report a network condition only once. The remaining areas that most benefit from a human look are the GPU ECC check (N/A on cards without ECC support) and the SSD Wear gauge (a low gauge on a lightly used drive is healthy, not a fault). When a finding is not obvious, open the Why? link for that check and confirm with the suggested command before treating the host as faulty.

References

  • HPCMATE Portal — Diagnostic view (POST/GET /api/servers/:id/diagnostic)
  • diagnose_server.py — the read-only check implementation
  • NVIDIA nvidia-smi documentation (temperature, ECC, XID)
  • smartmontools (smartctl) documentation
  • lm-sensors documentation

Related Pages