bench_press 0.4.1
bench_press: ^0.4.1 copied to clipboard
A modern, statistically sound, compiler-aware multi-runtime benchmarking framework for Dart and Flutter.
0.4.1 #
- Interleaved
BenchmarkGrouptrials now run in blocks (sized for fourABBA BAABvisits overtrials) and discard two batches at every switch between variants. The first batches after a switch carry state left by the previous variant; with every trial a switch,0.4.0roughly doubled the per-cell spread (robust_cv) on Wasm, and a single discard recovered only half of that. Two discards per switch restore the sequential baseline at about +20% wall-clock. See #71.
0.4.0 #
- Breaking (library exports):
package:bench_press/bench_press.dartnow exports only the benchmark-authoring (Benchmark,AsyncBenchmark,BenchmarkGroup,BenchmarkMatrix,BenchmarkVariant,BenchmarkConfig,Throughput,ByteThroughput,ElementThroughput,Blackhole), suite entrypoint (mainBenchmark*), and.report()result (BenchmarkResult,BenchmarkMetrics,CalibratedBatch,WarmupResult) APIs. Internal CLI, subprocess runner, compiler, statistical helper, and telemetry schema types that were previously re-exported in0.3.1are now internal tolib/src/. - Breaking (config): removed the
matrix.entrypointskey frombench_press.yaml. It never chose which file ran: each value re-ran the same discovered benchmark files under a made-upentrypointcoordinate, so the matrix table compared identical runs. Pass benchmark files as positional paths tobench_press runinstead. A leftoverentrypoints:key is ignored. - Breaking (behavioral):
ByteThroughput.formatRatenow scales byte rates by decimal1000(KB/s,MB/s,GB/s) instead of1024, matching hardware memory bandwidth, network I/O conventions, andElementThroughput. - Breaking (behavioral):
bench_press diffandbench_press run --diffnow pair cells by full key (name, target, and every coordinate, including group) instead of by name and target alone. A name reused across groups was previously compared against the first match in the other run. The Benchmark column now shows coordinates, for exampleencode (group=A). Cells present in only one run are listed in anUnmatchednote, including when nothing pairs. - Breaking (behavioral): an explicitly configured Dart SDK (
sdkmatrix axis inbench_press.yaml) is now authoritative and fails fast with a specific diagnostic explaining why the path was rejected instead of silently falling back todartonPATH. BenchmarkGroup.report,BenchmarkMatrix.report, andmainBenchmarkSuitenow interleave measurement trials across group variants inABBA BAABrounds after every variant has finishedsetup, warmup,warmupComplete, and batch calibration. Previously, all trials of variantAran before variantBstarted, so a linear thermal ramp or a short host contention window duringBproduced a false speedup or regression while both variants still passed the within-variantisStablegate. WhenmaxTrialsis set, lockstep rounds continue across all variants in the group while any variant's CV exceeds 5%.bench_press runnow honorsdefaults.outputinbench_press.yaml, both for saving results and for the filerun --difflooks up. An explicit--outputor--savestill takes precedence, andbenchmark_results.jsonremains the fallback. A leading~expands to the home directory. The value must be a non-empty string; anything else is now a config error, where it was previously ignored.bench_press reportandbench_press diffdo not read it; pass them the path.- Added
--pin-cpu <cpu-list>tobench_press runto pin benchmark subprocesses (VM JIT, AOT, Node.js, and D8) viataskset -con Linux. - Added Fieller confidence interval and
isStablespeedup gating (gate: trueby default, configurable via--[no-]gateinbench_press run,report, anddiff). Unstable or unbounded-CI comparisons render asunresolvedand are excluded from geometric mean rollups; a resolved comparison whose 95% CI contains1.00xrenders as➖ ⚪ Neutral. - Added Markdown report warning banners to flag benchmarks whose declared
Throughput.bytesrate exceeds physical memory bandwidth (100 GB/s) or whose latency remains invariant (0.9x–1.25x) across>=8xpayload size spreads for a shared benchmark name across groups. - Derived post-warmup calibration batch sizes directly from steady-state warmup
convergence latencies (recorded in
warmup.estimated_op_ns), and added a Markdown report footnote listing any cell whose trial median moved more than 25% from the warmup estimate that sized its batch. - Updated
bench_press runandbench_press validateto discover and execute all positional file and directory paths supplied on the command line rather than silently ignoring arguments after the first path. - Invoked
warmupComplete()lifecycle hook duringBenchmarkVariantexecution, forwardedthroughputandwarmupComplete()forBenchmarkandAsyncBenchmarkinstances run viamainBenchmark*, and preserved programmaticBenchmarkConfigvalues unless explicitly overridden by CLI flags. - Fixed Markdown suite reporting so suites mixing standalone benchmarks and
BenchmarkGroupvariants collate grouped variants into a single multi-row### Group: ...comparison table with the baseline ordered first and a geometric mean summary footer, and stopped wrapping Markdown tables inmdformat off/mdformat onHTML-comment guards.
0.3.1 #
- Hardened
bench_press runandbench_press validateto exit with non-zero exit code (ExitCode.software) when target builds fail compilation, executions crash, or benchmark targets produce zero results. - Fixed
_finishSuiteExecutioninRunCommandto propagate partial failure statuses across multi-target and multi-file Cartesian matrix executions while still preserving valid accumulated results. - Implemented value equality (
operator ==) and order-independenthashCodeonDartSdk. - Fixed CLI-specified
--d8-pathand--node-pathoverrides inRunCommandandValidateCommandto ensure user-provided binary paths instantiate fresh compilers and process runners. - Fixed
ValidateCommand._validateCoordinateto iterate across all resolved runtime targets when Cartesian matrix configurations omit runtime dimensions (matchingRunCommandbehavior).
0.3.0 #
- Breaking Change: Streamlined
BlackholeAPI to a single universalconsume(Object? value)method. Removed redundant specialized methods (consumeInt,consumeDouble,consumeBool,consumeString,consumeObject). - Hardened
Blackholecompiler barrier against optimizing compiler Dead Code Elimination:- Adopted 3-bit cyclic Gray-code ring buffer indexing
(
(index & 7) ^ ((index & 7) >> 1)) to disrupt compiler loop unrolling and vectorization without consecutive slot collisions. - Coupled slot position with element hashing in
Blackhole.drain()viaObject.hash(_sink[i], i)to guarantee position-dependent reduction. - Corrected documentation regarding retention guarantees (retains the last 8 writes across the cyclic Gray-code buffer).
- Fixed 5.2x latency cliff on Web/JavaScript previously caused by eager
double.hashCodecomputation.
- Adopted 3-bit cyclic Gray-code ring buffer indexing
(
- Added top-level
### Suite Summaryroll-up table with geometric mean, minimum, and maximum speedup across comparison groups toMarkdownReporter.renderSuiteandMarkdownReporter.renderSuiteSummaryTablewhen suites contain 2 or more distinct groups (Issue #22). - Added compilation artifact caching to
TargetCompiler.compileand--cache/--no-cacheCLI flags tobench_press run, avoiding redundant AOT/Wasm/JS compilation for unchanged benchmark files and dependencies (Issue #21). - Added explicit CLI options
--d8-pathand--node-pathtobench_press runandbench_press validate, supporting custom binary overrides,D8_PATHandNODE_BINARY/NODE_EXECUTABLEenvironment variables, and SDK auto-probing for bundled D8 underbin/resources/dart2wasm/d8(Issue #19). - Fixed silent crashes on Node.js for JS and Wasm targets by replacing
stdout.writelnwithprint, generating self-invoking.run.mjsand.node.cjswrappers with unhandled rejection listeners, and forwarding CLI arguments viadartMainRunner(Issue #29, #30). - Added parameterized matrix group builder
BenchmarkGroup.matrix<T>(and convenienceBenchmark.matrix<T>) andBenchmarkMatrix<T>to benchmark competing implementations across parameterized inputs or datasets without repetitive boilerplate (Issue #23). - Added
mainBenchmarkMatrixCLI entrypoint and updatedmainBenchmarkSuiteto executeBenchmarkMatrixinstances seamlessly. - Added robust dispersion metrics to
BenchmarkMetrics: Median Absolute Deviation (madNs), normal-consistent robust CV (robustCv = (1.4826 * madNs) / medianNs), and Interquartile Range (iqrNs). - Added
isRobustStabletoBenchmarkMetricsand updatedisStableto incorporate robust dispersion, preventing transient bimodal GC sweeps from falsely failing steady-state stability for allocation-heavy workloads (Issue #20). - Added adaptive trial scaling via
--max-trialsCLI option andmaxTrialsinBenchmarkConfig/DefaultsConfig, allowingBenchmarkRunnerto dynamically collect additional measurement trials when initial variance exceeds threshold. - Enhanced
AdaptiveWarmupDetectorto distinguish systemic monotonic drift from transient bimodal outliers viacomputeRobustSemandhasSystemicDrift. - Added non-breaking
warmupComplete()lifecycle hook toBenchmarkandAsyncBenchmark.
0.2.0 #
-
Added multi-tier Cartesian comparison matrix support (Issue #5) via unified
bench_press.yamlconfiguration manifest,--config, and--dry-runinspection flag. -
Added N-dimensional
coordinates: Map<String, String>mapping toBenchmarkEntrytelemetry schema, replacing the single-axisgroupproperty. -
Added
MarkdownReporter.renderMatrixComparisonTableto render multidimensional matrix reports with grouped left-hand dimension columns and Fieller 95% ratio confidence intervals. -
Extracted mathematical and calibration constants (
Lanczos,Acklam, andBenchmarkCalibratorthresholds) with detailed doc comments. -
Removed legacy transitional
--compare-sdkoption in favor of unified Cartesian matrix configurations inbench_press.yaml. -
Added positional argument support (
<baseline> [current]) tobench_press diffalongside--baseline(-b) and--current(-c). -
Added
Blackhole.consumeStringandBlackhole.consumeObjectoverloads with@pragma('dart2js:never-inline')compiler barriers. -
Added
maxSemRelativeErroroption (default0.03) toBenchmarkConfigfor steady-state warmup convergence. -
Added
BenchmarkEntry.copyWithmethod and updatedBenchmarkEntry.keyto include optionalgroup($name:$target:$group) with deterministic name/target ordering inBenchmarkSuiteResult.deepMerge. -
Updated
BenchmarkSuiteResult.groupsto return group names in order of appearance rather than sorted alphabetically. -
Aligned default
targetBatchDurationto100msacross CLI runners andBenchmarkConfig. -
Unified default benchmark discovery directories across
runandvalidatecommands (benchmark,benchmarks,bench). -
Simplified benchmark discovery to convention-based matching (targeting files ending in
*_benchmark.dartor*_bench.dart) requiring standardvoid main()entrypoints, eliminating ad-hoc regex content parsing and dynamic wrapper script generation. -
Updated benchmark discovery to default strictly to
benchmark/(orbenchmarks/,bench/) and throw explicit errors (FormatExceptionfor non-Dart files,PathNotFoundExceptionfor nonexistent paths) rather than silently ignoring files or walking the entire repository root. -
Removed
BenchmarkFileKindenum and dynamic wrapper script generation. -
Removed deprecated
KbssdWarmupDetectoralias in favor ofAdaptiveWarmupDetector. -
Fixed steady-state warmup convergence math using Standard Error of the Mean (SEM) relative error (
<= 3%) and stationarity checks. -
Hardened
Blackhole.drain()compiler barrier against whole-program Dead-Store Elimination across AOT, Wasm, and JavaScript. -
Fixed
BenchmarkCalibratorto support sub-10µs operations without throwingCalibrationException, while throwingCalibrationExceptionby default when maximum probe batches produce zero elapsed ticks (elapsedUs == 0) unlessforceRun: true(--force-run) is specified (which warns and continues). -
Added
mode('sync'vs'async') property toBenchmarkResult(whichBenchmarkEntry.fromResultnow inherits for JSON telemetry). -
Implemented continuous Student's t-distribution quantile calculation (regularized incomplete beta for
1 < df < 2and Hill's Algorithm 396 fordf > 2) for accurate Fieller confidence intervals across all degrees of freedom. -
Fixed unhandled exception propagation in Isolate execution mode (
BenchmarkProcessRunner). -
Prevented floating-point overflow in geometric mean speedup reporting via log-sum calculation.
-
Updated
FiellerInterval.computeto returnisValid: false(withNaNbounds) when sample size is degenerate (N < 2).
0.1.0 #
- Initial release of
bench_press: A modern, statistically sound, compiler-aware multi-runtime benchmarking framework for Dart and Flutter. - Multi-runtime execution support across JIT, AOT (
dart compile exe), WasmGC (dart compile wasm), and JavaScript (dart compile js). Benchmark,AsyncBenchmark,BenchmarkVariant, andBenchmarkGroupharnesses with lifecycle hooks (setup,run,teardown).Blackholedead-code elimination (DCE) sink to safely consume benchmark results without compiler dead-code stripping.Throughputmetric tracking for byte rates (B/s,KB/s,MB/s,GB/s) and element rates (items/s,records/s,tokens/s).- Automated batch calibration and steady-state warmup convergence detection.
- Statistical summary metrics (Mean, Median, Min, Max, StdDev, CV, p95, p99, Ops/sec) with Fieller 95% confidence intervals for variant ratios.
- Markdown reporting with side-by-side variant comparisons and before/after baseline diffing.
bench_pressCLI withrun,validate,report, anddiffsubcommands.