Core

Metrics

The native engine records every request-like step into a fixed-size HDR histogram and folds custom counters/histograms from action outputs into the same run aggregates. Two outputs: a machine-readable [stats] line every 5 seconds during the run, and a k6-compatible summary at the end.

Request metrics (http_req_*)

Actions that perform a timed request return an HttpSample (duration_ms, status, failed). Durations are recorded in microseconds — sub-millisecond loopback calls stay distinguishable.

ActionWhat feeds http_req_durationWhat counts as failed
std/http@v1Every requestStatus ≥ 400, transport error, timeout (logged as → TIMEOUT after …ms; status reported as 0)
std/tcp@v1Connect + send/read exchangeConnect failure, timeout, expect mismatch
std/udp@v1Send (+ optional reply wait)Timeout, expect mismatch
std/ws@v1The whole one-shot session (one sample)Handshake/transport error, an until_* rule not met in time
std/ws-connect@v1The handshake onlyHandshake failure (connected: false)

Deliberately not feeding http_req_duration — a step whose duration says nothing about target latency would poison the shared percentiles:

ActionWhy
std/ws-recv@v1How long a server waits before pushing is not target latency
std/ws-ping@v1Transport RTT; bound it with check: { duration_ms_lt: … } instead
std/ws-send@v1, std/ws-close@v1No meaningful request latency
std/grpc*@v1 (whole family)gRPC has its own histograms (grpc_req_duration, grpc_msg_rtt); stream lifetimes span user steps, so streams don't feed even those
std/db-*@v1 (whole family)DB steps have their own histograms (db_connect_duration, db_query_duration) and counters (db_rows, db_errors)
std/child_process@v1, std/kill_process@v1Process lifecycle, not requests
std/check@v1, std/sleep@v1, std/log@v1, std/file-*@v1No network I/O

Note the asymmetry: gRPC assertion failures go to grpc_req_failed (a custom counter, driven by expect_status), not to http_req_failed.

Final summary

Printed at the end of every run, k6-compatible so downstream parsers (dashboards, perfscale serve) treat all engines uniformly:

vus....................: 10 min=1 max=10
iterations..............: 4521 150.23/s
http_req_duration......: avg=0.42ms p(50)=0.31ms p(90)=0.88ms p(95)=1.02ms p(99)=1.90ms min=0.09ms max=3.10ms
http_req_failed........: 0.00%
http_reqs..............: 4521 150.23/s
  • vus / iterations (+ per-second rate) are always emitted, even for sleep-only runs.
  • dropped_iterations is always emitted for arrival-rate runs — 0 when the pool kept up — so threshold gates (dropped_iterations: ["count==0"]) resolve instead of erroring on an unknown metric.
  • The http_req_* block appears only when at least one sample was recorded.
  • Percentiles come from a fixed-size HDR histogram (1 µs – 1 h range, two significant figures → ≤1% quantile error) — memory stays flat no matter how long the soak, at the cost of an error invisible at the printed precision.

Live [stats] lines

Every 5 seconds while VUs run, one machine-readable line:

[stats] ts=1720000000000 rps=246.80 err_pct=0.00 p50=1.20 p90=3.40 p95=4.10 p99=8.20 reqs=1234 iters=456
  • ts — unix epoch milliseconds; rps — requests in the just-finished 5 s window; err_pct — cumulative failure percentage; p50p99 — cumulative percentiles in ms (the histogram is never reset, so they converge instead of jittering); reqs — cumulative requests; iters — cumulative iterations.
  • With no requests yet the percentiles are omitted: [stats] ts=… rps=0.00 reqs=0 iters=3.
  • These lines exist for streaming consumers (the controlplane parses them out of the log stream); see --quiet for console behavior.

Custom metrics (value.metrics)

Any action can attach a reserved metrics object to its step output; the runner folds it into the run aggregates:

  • A number becomes a counter, summed across VUs and iterations, reported as <name>: <total> <rate>/s.
  • An array of numbers becomes HDR histogram samples (milliseconds), reported as <name>: avg=…ms p(50)=… p(90)=… p(95)=… p(99)=… min=… max=… count=N.

Built-in emitters:

NameTypeEmitted byMeaning
ws_msgs_sentcounterstd/ws@v1, std/ws-send@v1WS messages sent
ws_msgs_receivedcounterstd/ws@v1, std/ws-recv@v1WS messages read
ws_msg_rtthistogramstd/ws@v1, std/ws-recv@v1Send → first matching reply (application-level RTT)
pubsub_msgs_publishedcounterstd/pubsub@v1Messages accepted by the transport
pubsub_msgs_receivedcounterstd/pubsub@v1 (with subscribe)Messages counted toward subscribe.count
pubsub_e2e_mshistogramstd/pubsub@v1 (with subscribe)Publish-phase start → message consumed, one sample per matched message
shared_variable_wait_mshistogramstd/get_shared_variable@v1 (with wait_for)Time blocked in wait_for before the condition held, one sample per waiting read
llm_ttft_mshistogramstd/llm@v1 (streamed)Request start → first content chunk (time to first token)
llm_tokens_per_sechistogramstd/llm@v1Completion tokens / generation time (after the first token when streamed)
llm_prompt_tokenscounterstd/llm@v1Prompt tokens as reported by the server
llm_completion_tokenscounterstd/llm@v1Completion tokens as reported by the server
llm_chunkscounterstd/llm@v1 (streamed)SSE chunks received
grpc_req_durationhistogramstd/grpc@v1, std/grpc-call@v1Unary call latency
graphql_req_durationhistogramstd/graphql@v1GraphQL operation round trip
graphql_errorscounterstd/graphql@v1GraphQL-level errors, including partial-data responses that pass the step
graphql_op_<operationName>_durationhistogramstd/graphql@v1Per-operation latency, only for named operations (bounded cardinality)
grpc_msg_rtthistogramstd/grpc-call@v1, std/grpc-stream-recv@v1Send → matching reply RTT
grpc_msgs_sentcounterstd/grpc-call@v1, std/grpc-stream-send@v1gRPC messages sent
grpc_msgs_receivedcounterstd/grpc-call@v1, std/grpc-stream-recv@v1, std/grpc-stream-close@v1gRPC messages read
grpc_req_failedcounterstd/grpc-call@v1, std/grpc-stream-close@v1Calls that missed expect_status
db_connect_durationhistogramstd/db-connect@v1 (success)Connect + pool setup latency
db_query_durationhistogramstd/db-query@v1, std/db-tx-*@v1Query latency; includes the fresh connect in per-query mode
db_rowscounterstd/db-query@v1Rows returned, or rows affected when the statement returned none
db_errorscounterstd/db-*@v1 (failure)Failed DB steps, total. Successful DB steps emit db_errors: 0, so the counter exists (at 0) on fully healthy runs — gates like db_errors: ["count==0"] work either way
db_errors_connection / _constraint / _deadlock / _timeout / _othercounterstd/db-*@v1 (failure)Same, split by class (SQLSTATE / errno / SQLite result code)

Downstream actions use the same channel — e.g. the proprietary FIX action emits fix_messages_sent.

Failure-rate metrics (<family>_failed)

Alongside the metrics payload, the runner derives per-invocation failure samples generically: for every histogram (array-valued) metric an invocation emits, it records one 0/1 sample — 1 when the step invocation failed, 0 when it succeeded — under the metric's family name with a trailing _duration/_rtt replaced by _failed:

Duration metricDerived failure metric
http_req_durationhttp_req_failed (native to the HTTP path)
db_query_durationdb_query_failed
db_connect_durationdb_connect_failed
grpc_req_durationgrpc_req_failed
graphql_req_durationgraphql_req_failed
ws_msg_rttws_msg_failed
pubsub_e2e_mspubsub_e2e_ms_failed
llm_ttft_msllm_ttft_ms_failed
llm_tokens_per_secllm_tokens_per_sec_failed

These print as <name>: <pct>% (k6's http_req_failed shape). Because one sample is recorded per invocation (not per duration sample), failed/total over them is exactly the step family's failure rate — that is what std/thresholds@v1 evaluates with rate, e.g. db_query_failed: ["rate<0.05"]. Note a failed step that emits no duration sample (e.g. db-connect that never connected) records no sample either, so its family rate covers completed invocations.

When a family already has a same-named counter (the gRPC actions emit a grpc_req_failed counter for expect_status misses), the rate metric shadows it in the summary and in threshold evaluation.

Run-level gates (std/thresholds@v1)

A std/thresholds@v1 step (typically in after:) evaluates k6-style expressions against the run aggregates and prints one machine-readable line after the metric summary:

thresholds: {"status":"fail","message":"db_query_failed rate=1 ≥ 0.05; checkout SLO","violations":[{"metric":"db_query_failed","expr":"rate<0.05","actual":1.0}]}

Aggregates come from the same HDR histograms/counters as the text summary, so gate numbers match what the summary prints. The line is collected into perfscale run --summary-export output under thresholds ({status, message, violations}), and a fail status makes the CLI exit non-zero. See actions.md.

GPU metrics (gpu:)

With gpu.enabled: true in the run config, the native engine samples every GPU on the host — utilization %, VRAM used/total, temperature, power draw — once per interval_ms for the whole VU phase (via nvidia-smi or a dcgm-exporter endpoint). After the metric summary the run prints a compact per-device block plus one machine-readable gpu: {...} line with the full timeseries, which --summary-export embeds under gpu:

gpu: 1 device, 300 samples every 1000ms (nvidia-smi)
gpu0: util avg=64.3% max=100.0% vram max=41088/81559MiB temp max=71.0C power max=512.3W
gpu: {"source":"nvidia-smi","interval_ms":1000,"devices":[{"index":0,"samples":[…],…}]}

Collection is best-effort: no GPU / missing tooling logs one warning and the run continues without GPU metrics. Sample timestamps share the epoch-ms timeline with the [stats] lines, so load and GPU state correlate directly — the primary use case is std/llm@v1 runs against local model servers. Full guide: gpu.md.

--quiet

Two independent layers:

  • At the source (native engine): per-iteration success output — request lines, sleep markers, passing checks — is not even formatted or sent. Errors, failing checks, [stats] lines, and the final summary are always emitted into the stream.
  • At the CLI printer: under --quiet, stdout lines print only if they are k6-shaped summary lines (vus, iterations, http_req_*); stderr and system lines always print. Custom metric lines and [stats] stay in the stream for log consumers but are hidden from the console.

Forwarding the summary (report)

Point the run at a perfscale serve instance and the summary lines are forwarded when the run finishes:

# config.yaml
report:
  url: http://localhost:7999
perfscale run -f test.yaml -c config.yaml --report http://localhost:7999

The CLI flag wins over the config block. After the run, the CLI POSTs the collected summary lines (only the k6-shaped ones) as {"lines": […]} to <url>/api/v1/metrics with a 5 s timeout; delivery problems are logged as [report] … on stderr and never fail the run itself.