nupp.bench

Measuring Nupp code from inside an ordinary program.

A benchmark here is a program, not a case a runner discovered. It links this module, runs under nupp run, and reports itself:

local bench = nupp.bench

local function sized(b: nupp.bench.Case): nil
    for _ = 1, b.n do
        bench.keep(makePoint())
    end
end

bench.case("presize.sized", sized)
bench.report()

The library and nupp bench have different jobs. An application's hot loop lives in the application, so only something the loop can call can measure it. The command discovers those programs and gives each named case an isolated process.

report raises rather than returning a status, because a chunk's return value is discarded and a run that did not raise exits zero. It writes the record first, so a case that trips the gate is still a case whose record the runner can merge.

Module contents

Types

TypeKindDescription
CaserecordOne case handed to a measured body.
FormatOptionstypeOptional sections in the human result report.
FrameSessionrecordA running frame measurement, as frames returns it.
InvocationrecordOne expanded suite case handed to setup, run and teardown callbacks.
MeasurementrecordWhat one measured case contributes to a record.
SuiteCasetypeOne workload in a suite.
SuiteOptionstypeA comparative benchmark suite.
VarianttypeOne implementation of every case in a suite.

Functions

FunctionKindDescription
casefunctionMeasures one case and adds it to the record this program will report.
comparefunctionCompares one record against one baseline record and returns what moved.
decodeBaselinefunctionDecodes a baseline record, refusing one this version cannot read.
formatfunctionFormats a record as the compact result table used by both a standalone program and the set runner.
framesfunctionOpens a frame measurement.
keepfunctionKeeps a value the measured body produced, so the work that produced it survives.
reportfunctionWrites the record, compares it against a baseline when one was named, and raises when a gated counter moved.
suitefunctionDeclares a comparative suite.

Types#

Caserecord#

record bench.Case ...

One case handed to a measured body.

Members

NameKindDescription
namefieldThe case's name, as case was given it.
nfieldIterations the body is to run.

namefield#

name: string

The case's name, as case was given it.

nfield#

n: integer

Iterations the body is to run. Fixed before the measured rounds begin, and carried in the record, so a comparison can re-run the same work.

FormatOptionstype#

type bench.FormatOptions = {
    --- Compare suite variants by their geometric-mean ratio.
    geometricMean: boolean?,

    --- Forks behind the measurements, when the runner merged them.
    ---
    --- Decides which columns the table carries. One fork reports a within-process
    --- interquartile range, which is a spread and not an interval; more than one
    --- reports the interval and the coverage it attained.
    forks: integer?,

    --- Lines appended below the table, each already wrapped.
    notes: {string}?
}

Optional sections in the human result report.

FrameSessionrecord#

record bench.FrameSession ...

A running frame measurement, as frames returns it.

The application owns its loop, so this offers a condition and a report rather than taking the loop over. There is no exit-time fallback: nothing hands this module a callback when the chunk returns, collection before shutdown is not guaranteed, and the compiler's entry point ends in os.exit. A session that is never reported produces no record, and the runner says so.

Members

NameKindDescription
namefield
budgetMsfield
wantedfield
samplesfield
startedfield
sessionfield
samplingfield
donefield
moremethodWhether more frames are wanted.
beginmethodMarks the start of a frame.
finishmethodMarks the end of a frame and records it.
reportmethodEnds the session, records it, and reports.

namefield#

name: string

budgetMsfield#

budgetMs: number

wantedfield#

wanted: integer

samplesfield#

samples: {number}

startedfield#

started: number

sessionfield#

session: any

samplingfield#

sampling: any

donefield#

done: boolean

moremethod#

more: function bench.FrameSession:more(): boolean

Whether more frames are wanted. False once count frames have been recorded, and always true when the session was opened without one.

Returns
TypeDescription
boolean

beginmethod#

begin: function bench.FrameSession:begin(): nil

Marks the start of a frame.

Returns
TypeDescription
nil

finishmethod#

finish: function bench.FrameSession:finish(): nil

Marks the end of a frame and records it.

Elapsed on the monotonic clock, not os.clock, which counts processor time. A frame that waited on presentation, on I/O, or on a sleep spends very little of either and blows its budget anyway, and missing the budget is the whole measurement.

Returns
TypeDescription
nil

reportmethod#

report: function bench.FrameSession:report(): nil

Ends the session, records it, and reports.

A frame report is a distribution rather than a median: a budget missed one frame in a hundred is a visible stutter and an unmoved mean.

This writes the record and applies the gate, which is why an application's loop needs no second call. bench.report is idempotent, so a program with both cases and frames may still call it itself.

Returns
TypeDescription
nil
Raises
TypeCondition
string

when a gated counter moved, or the record cannot be written

Invocationrecord#

record bench.Invocation ...

One expanded suite case handed to setup, run and teardown callbacks.

Members

NameKindDescription
namefieldThe unqualified case name from the suite declaration.
datafieldCase-specific data supplied by the declaration.
parametersfieldOne value for each parameter in the expanded case.

namefield#

name: string

The unqualified case name from the suite declaration.

datafield#

data: any

Case-specific data supplied by the declaration.

parametersfield#

parameters: {[string]: any}

One value for each parameter in the expanded case.

Measurementrecord#

record bench.Measurement ...

What one measured case contributes to a record.

Members

NameKindDescription
namefieldThe case's name.
kindfieldcase, suite or frames.
nfieldIterations per round, for a case.
roundsfieldRounds timed, for a case.
medianMsfieldThe median round, in milliseconds.
minMsfieldThe fastest round, in milliseconds.
samplesMsfieldSuite samples in milliseconds, normalized to one represented operation.
meanMsfield
stdevMsfield
p90Msfield
p99MsfieldThe 99th percentile of this measurement's samples, in milliseconds: rounds for a case, samples for a suite, frames...
sampleIterationsfieldCalls to a suite variant inside one timed sample, and the operations one call represents.
operationsPerInvocationfield
warmupIterationsfield
minSamplesfield
minDurationMsfield
maxSamplesfield
maxDurationMsfield
suitefieldComparative-suite identity.
variantfield
caseNamefield
parametersfield
baselineVariantfield
allocatedKbfieldKilobytes allocated, read with the collector stopped.
retainedKbfieldKilobytes still held once the collector has run again, which is what "retained" has to mean.
p50MsfieldFrame durations in milliseconds at the named quantiles.
p999Msfield
framesfieldFrames recorded, and how many exceeded the budget.
overBudgetfield
abortSitesfieldAbort site identities seen while measuring, each severity|reason|location|zone.
totalAbortsfieldAbort events and blacklistings behind those sites.
blacklistedfield
profilePathfieldCollapsed-stack profile written for this measured window, when requested.
p25MsfieldQuantiles of this measurement's own samples, in milliseconds.
p75Msfield
iqrMsfieldThe interquartile range of this measurement's own samples, in milliseconds.
trendfieldWhat a monotone-trend test made of this process's samples in execution order: trend, no-trend-detected, or unknown.
trendTaufieldKendall's tau and the two-sided p-value behind trend, and how far the level actually moved between the ends of the...
trendPValuefield
trendDriftfield
concentrationfieldThe fraction of this measurement's samples lying within ten percent of its median.
outlierCountfieldSamples beyond three interquartile ranges of the nearer quartile, and the largest of them over the median.
outlierMaxRatiofield
forkCountfieldProcesses this measurement summarizes, when the runner merged forks into it.
forkSummariesMsfieldOne normalized summary per fork, in execution order.
intervalLowMsfieldThe interval for this benchmark's population median, when one is available.
intervalHighMsfield
intervalCoveragefield
intervalWithheldfieldWhy there is no interval: below-minimum-forks or trend-warning.

namefield#

name: string

The case's name.

kindfield#

kind: string

case, suite or frames.

nfield#

n: integer?

Iterations per round, for a case.

roundsfield#

rounds: integer?

Rounds timed, for a case.

medianMsfield#

medianMs: number?

The median round, in milliseconds. Recorded, never gated.

minMsfield#

minMs: number?

The fastest round, in milliseconds.

samplesMsfield#

samplesMs: {number}?

Suite samples in milliseconds, normalized to one represented operation.

meanMsfield#

meanMs: number?

stdevMsfield#

stdevMs: number?

p90Msfield#

p90Ms: number?

p99Msfield#

p99Ms: number?

The 99th percentile of this measurement's samples, in milliseconds: rounds for a case, samples for a suite, frames for a frame session.

sampleIterationsfield#

sampleIterations: integer?

Calls to a suite variant inside one timed sample, and the operations one call represents.

operationsPerInvocationfield#

operationsPerInvocation: integer?

warmupIterationsfield#

warmupIterations: integer?

minSamplesfield#

minSamples: integer?

minDurationMsfield#

minDurationMs: number?

maxSamplesfield#

maxSamples: integer?

maxDurationMsfield#

maxDurationMs: number?

suitefield#

suite: string?

Comparative-suite identity. Kept apart so a reader need not parse name.

variantfield#

variant: string?

caseNamefield#

caseName: string?

parametersfield#

parameters: {[string]: any}?

baselineVariantfield#

baselineVariant: string?

allocatedKbfield#

allocatedKb: number?

Kilobytes allocated, read with the collector stopped. A case stores one round; a suite stores one represented operation. A size rather than a count of allocations, and recorded only.

retainedKbfield#

retainedKb: number?

Kilobytes still held once the collector has run again, which is what "retained" has to mean. An earlier implementation reported the number above and called it this one: with collection stopped for the rounds, the heap delta is what they allocated, not what survived them.

p50Msfield#

p50Ms: number?

Frame durations in milliseconds at the named quantiles. p99Ms is above.

p999Msfield#

p999Ms: number?

framesfield#

frames: integer?

Frames recorded, and how many exceeded the budget.

overBudgetfield#

overBudget: integer?

abortSitesfield#

abortSites: {string}?

Abort site identities seen while measuring, each severity|reason|location|zone. Identities rather than counts: a count is partly a function of how much work ran, and a loop that started aborting is the finding at any count.

Nil when no session could be opened, which is not the same as an empty set: one says nothing aborted and the other says nobody looked.

totalAbortsfield#

totalAborts: integer?

Abort events and blacklistings behind those sites. Recorded, never gated.

blacklistedfield#

blacklisted: integer?

profilePathfield#

profilePath: string?

Collapsed-stack profile written for this measured window, when requested.

p25Msfield#

p25Ms: number?

Quantiles of this measurement's own samples, in milliseconds.

The table shows [p25, p99] beside the score, which is a range rather than a percentage either side of the median on purpose: a benchmark whose samples are bimodal puts its median in the gap between the two clusters, and a symmetric "+-x%" around that median describes a distribution the run never produced.

The upper end is the tail rather than the third quartile because the tail is what a reader is looking for. A slow mode holding a tenth of the samples moves p99 and leaves p75 where it was, so a box would have hidden exactly the case worth seeing. p75Ms stays in the record for anyone who wants the box.

p75Msfield#

p75Ms: number?

iqrMsfield#

iqrMs: number?

The interquartile range of this measurement's own samples, in milliseconds.

Within-process spread. It says how much the samples inside one process varied and is not a confidence interval: those samples share a heap, a set of compiled traces and a thermal state, so they are not independent draws and no interval may be derived from them. An interval comes from fork summaries, which the runner assembles across processes.

trendfield#

trend: string?

What a monotone-trend test made of this process's samples in execution order: trend, no-trend-detected, or unknown.

There is no value asserting a steady state. A test that found no trend has not established one, and saying otherwise is the mistake this field exists to avoid making.

trendTaufield#

trendTau: number?

Kendall's tau and the two-sided p-value behind trend, and how far the level actually moved between the ends of the series.

trend needs both a significant test and a drift worth acting on, so a series that moved half a percent with a p-value of 0.005 is not reported as trending.

trendPValuefield#

trendPValue: number?

trendDriftfield#

trendDrift: number?

concentrationfield#

concentration: number?

The fraction of this measurement's samples lying within ten percent of its median.

Near one when the samples describe a single rate. Near zero when they split into clusters and the median falls in the gap, which is what a collector running on alternate samples produces and which makes the reported score describe a rate the benchmark never ran at.

outlierCountfield#

outlierCount: integer?

Samples beyond three interquartile ranges of the nearer quartile, and the largest of them over the median.

Classified, never removed. A real warmup or deoptimization phase falls exactly where these fences do, so excluding what they catch would delete the behavior trend is looking for. Every sample stays in samplesMs and in the median.

outlierMaxRatiofield#

outlierMaxRatio: number?

forkCountfield#

forkCount: integer?

Processes this measurement summarizes, when the runner merged forks into it.

Absent on a measurement a child wrote: a child is one process and knows nothing about the others.

forkSummariesMsfield#

forkSummariesMs: {number}?

One normalized summary per fork, in execution order. The input the interval below was computed from, kept so a reader can recompute it.

intervalLowMsfield#

intervalLowMs: number?

The interval for this benchmark's population median, when one is available.

intervalCoverage is the probability this interval actually attains, not one that was requested: it falls out of the fork count through the sign test, and a run with too few forks gets no interval rather than a relabelled one.

intervalHighMsfield#

intervalHighMs: number?

intervalCoveragefield#

intervalCoverage: number?

intervalWithheldfield#

intervalWithheld: string?

Why there is no interval: below-minimum-forks or trend-warning.

SuiteCasetype#

type bench.SuiteCase = {
    name: string,
    data: any?,
    parameters: {[string]: {any}}?,

    --- Operations represented by one call to a variant's `run` callback.
    operations: integer?
}

One workload in a suite. Parameter lists are expanded as a Cartesian product.

SuiteOptionstype#

type bench.SuiteOptions = {
    name: string,
    variants: {bench.Variant},
    cases: {bench.SuiteCase},

    --- Variant used as the denominator in the human result table.
    baselineVariant: string?,

    --- Calls made before timing begins.
    warmupIterations: integer?,

    --- Calls to `run` inside one timed sample. State is set up once per sample.
    sampleIterations: integer?,

    --- Stop once both this many samples and this many milliseconds of measured time
    --- exist.
    minSamples: integer?,
    minDurationMs: number?,

    --- Safety bounds when a sample is much faster or slower than expected.
    maxSamples: integer?,
    maxDurationMs: number?
}

A comparative benchmark suite.

Varianttype#

type bench.Variant = {
    name: string,
    setup: (function(bench.Invocation): any)?,
    run: function(any, bench.Invocation): any,
    teardown: (function(any, bench.Invocation): nil)?
}

One implementation of every case in a suite.

Functions#

bench.casefunction#

function bench.case(name: string, body: function(bench.Case): nil, fixedN: integer?): nil

Measures one case and adds it to the record this program will report.

The body is handed the case and runs its own loop. Calling a one-iteration closure n times instead would put a call boundary inside the measurement and change what the recorder sees, so the loop belongs to the body.

Arguments

NameTypeDescription
namestring

what to record the case as

bodyfunction(bench.Case): nil

the work, which iterates case.n times

fixedNinteger?

an explicit iteration count, skipping calibration

Returns

TypeDescription
nil

Raises

TypeCondition
string

when the name is empty or already declared

bench.comparefunction#

function bench.compare(record: any, baseline: any): {string}, {string}

Compares one record against one baseline record and returns what moved.

Exported because the runner owns the baseline for a whole set and has to apply the same rules a standalone case applies to its own. Two copies of these rules would disagree, and the one that mattered would be whichever the reader was not looking at.

A comparison with no comparable baseline is reported as exactly that. It is a result, and a different one from a regression: the runner has to say which, or a lost baseline reads like a pass.

Arguments

NameTypeDescription
recordany

what this run measured

baselineany

the record to compare it against

Returns

TypeDescription
{string}

the gated differences, which are failures

{string}

the differences that were only reported

bench.decodeBaselinefunction#

function bench.decodeBaseline(text: string, path: string, oldest: integer?): any, string?

Decodes a baseline record, refusing one this version cannot read.

Anything that decodes as JSON used to read as a baseline: a record from a newer bench, or a file that is not a record at all, matched no case, so every case read as having no baseline, and that looks like a pass. Exported so the runner applies the same rule to its own documents, from the oldest schema it reads.

Arguments

NameTypeDescription
textstring

what the baseline file holds

pathstring

where it was read from, for the message

oldestinteger?

the oldest schema the reader accepts, OLDEST_BASELINE_SCHEMA by default

Returns

TypeDescription
any

the decoded baseline, or nil where it is refused

string?

why it was refused, naming its schema and the ones this version reads

bench.formatfunction#

function bench.format(record: any, options: bench.FormatOptions?): string

Formats a record as the compact result table used by both a standalone program and the set runner. The shape follows the useful part of JMH's final report, but names the statistic p50: these scores are medians, not averages with confidence intervals.

Case scores are normalized to one body iteration. The calibrated batch size stays in the JSON record, where the runner also finds it for the next comparison.

Arguments

NameTypeDescription
recordany
optionsbench.FormatOptions?

Returns

TypeDescription
string

bench.framesfunction#

function bench.frames(name: string, budgetMs: number?, count: integer?): bench.FrameSession

Opens a frame measurement.

Plain parameters rather than an options record, because a record would make the ordinary call site write new bench.FrameOptions(...) -- a table literal is not one, which is what the first version of this API got wrong in its documented example.

Arguments

NameTypeDescription
namestring

what to record the session as

budgetMsnumber?

milliseconds a frame is allowed; frames over it are counted

countinteger?

frames to record before more turns false, or nil to let the application decide when to stop

Returns

TypeDescription
bench.FrameSession

Raises

TypeCondition
string

when the name is empty or already declared

bench.keepfunction#

function bench.keep(value: any): nil

Keeps a value the measured body produced, so the work that produced it survives.

This is the one rule writing a case requires. LuaJIT removes work whose result does not escape its trace, which is correct and is also the most common way a benchmark comes out impossibly fast: the loop under test is deleted and the measurement is of nothing.

The store is to a field of this module, which is visible outside any trace and so cannot be sunk. Nothing checks that this stays true of LuaJIT; a collapse in the reported durations is the only signal, which is why it is worth reading them.

Arguments

NameTypeDescription
valueany

whatever the measured body produced

Returns

TypeDescription
nil

bench.reportfunction#

function bench.report(): nil

Writes the record, compares it against a baseline when one was named, and raises when a gated counter moved.

The write happens first on purpose. A case that trips the gate is a case whose record the runner has to merge, or the report says only that something failed.

Raising is how a status reaches the shell: a chunk's return value is discarded and a run that did not raise exits zero. os.exit would also produce one and would discard the run's own profile and trace-abort reports, which are written after the chunk returns.

Returns

TypeDescription
nil

Raises

TypeCondition
string

when a gated counter moved, or the record cannot be written

bench.suitefunction#

function bench.suite(options: bench.SuiteOptions): nil

Declares a comparative suite. Every case, parameter set and variant is a named benchmark and therefore gets its own process under the set runner.

Arguments

NameTypeDescription
optionsbench.SuiteOptions

Returns

TypeDescription
nil

Raises

TypeCondition
string

when the declaration or its sampling bounds are invalid