Skip to main content

AcceleratorMetrics

One accelerator's live readings, normalized across collectors.

The collector-agnostic GPU/accelerator expression: any platform's collector (mactop on Apple Silicon, rocm-smi/sysfs on AMD, nvidia-smi on CUDA) fills the same shape, so the planner and dashboard reason about a heterogeneous fleet uniformly. A field a given collector cannot measure stays None (never a fake zero), so a reader can tell "0%" apart from "not reported". Units are fixed here so collectors normalize at their boundary: utilization_ratio is a 0..1 fraction, power is watts, temperature is degrees Celsius.

vendorVendor (string)

Possible values: [apple, amd, nvidia, intel, cpu, unknown]

Default value: unknown
nameName (string)
Default value: Unknown
utilizationRatio object
anyOf
number
vramTotalBytes object
anyOf
integer
vramUsedBytes object
anyOf
integer
gttTotalBytes object

GPU-mappable host (GTT) memory, for unified-memory APUs (e.g. AMD Strix Halo). On such a node the GPU addresses system RAM beyond the BIOS VRAM carve-out through GTT, so the usable GPU pool is far larger than vram_total_bytes (placement uses this to admit big models on a UMA node). None on discrete GPUs / collectors that do not report it.

anyOf
integer
powerWatts object
anyOf
number
temperatureCelsius object
anyOf
number
clockMhz object
anyOf
integer
computeCapability object

Discrete-GPU compute capability as "<major>.<minor>" (NVIDIA SM level), e.g. "8.0" (A100 Ampere), "9.0" (H100 Hopper), "10.0" (B100/B200 Blackwell), "12.0" (RTX 50 Blackwell). The engine/quant/placement decision keys on this, not on vendor: the same model+engine performs oppositely across generations (a benchmark showed vLLM's MXFP4 path losing single-stream on Ampere, which has no native FP4, but winning on Blackwell, which does). None on collectors that do not report it (AMD sysfs, Apple).

anyOf
string
nativeFp4 object

Whether the GPU accelerates FP4 natively (Blackwell sm100+, i.e. SM level (major, minor) >= (10, 0)). Derived at the collector boundary from the parsed compute capability (a numeric tuple compare, not a string compare). None when unmeasured.

anyOf
boolean
nativeFp8 object

Whether the GPU accelerates FP8 natively (Ada sm89 / Hopper sm90 and later). Derived at the collector boundary from the compute capability. None when unmeasured.

anyOf
boolean
AcceleratorMetrics
{
"vendor": "unknown",
"name": "Unknown",
"utilizationRatio": 0,
"vramTotalBytes": 0,
"vramUsedBytes": 0,
"gttTotalBytes": 0,
"powerWatts": 0,
"temperatureCelsius": 0,
"clockMhz": 0,
"computeCapability": "string",
"nativeFp4": true,
"nativeFp8": true
}