bench_press 0.3.1
bench_press: ^0.3.1 copied to clipboard
A modern, statistically sound, compiler-aware multi-runtime benchmarking framework for Dart and Flutter.
0.3.1 #
- Hardened
bench_press runandbench_press validateto exit with non-zero exit code (ExitCode.software) when target builds fail compilation, executions crash, or benchmark targets produce zero results. - Fixed
_finishSuiteExecutioninRunCommandto propagate partial failure statuses across multi-target and multi-file Cartesian matrix executions while still preserving valid accumulated results. - Implemented value equality (
operator ==) and order-independenthashCodeonDartSdk. - Fixed CLI-specified
--d8-pathand--node-pathoverrides inRunCommandandValidateCommandto ensure user-provided binary paths instantiate fresh compilers and process runners. - Fixed
ValidateCommand._validateCoordinateto iterate across all resolved runtime targets when Cartesian matrix configurations omit runtime dimensions (matchingRunCommandbehavior).
0.3.0 #
- Breaking Change: Streamlined
BlackholeAPI to a single universalconsume(Object? value)method. Removed redundant specialized methods (consumeInt,consumeDouble,consumeBool,consumeString,consumeObject). - Hardened
Blackholecompiler barrier against optimizing compiler Dead Code Elimination:- Adopted 3-bit cyclic Gray-code ring buffer indexing (
(index & 7) ^ ((index & 7) >> 1)) to disrupt compiler loop unrolling and vectorization without consecutive slot collisions. - Coupled slot position with element hashing in
Blackhole.drain()viaObject.hash(_sink[i], i)to guarantee position-dependent reduction. - Corrected documentation regarding retention guarantees (retains the last 8 writes across the cyclic Gray-code buffer).
- Fixed 5.2x latency cliff on Web/JavaScript previously caused by eager
double.hashCodecomputation.
- Adopted 3-bit cyclic Gray-code ring buffer indexing (
- Added top-level
### Suite Summaryroll-up table with geometric mean, minimum, and maximum speedup across comparison groups toMarkdownReporter.renderSuiteandMarkdownReporter.renderSuiteSummaryTablewhen suites contain 2 or more distinct groups (Issue #22). - Added compilation artifact caching to
TargetCompiler.compileand--cache/--no-cacheCLI flags tobench_press run, avoiding redundant AOT/Wasm/JS compilation for unchanged benchmark files and dependencies (Issue #21). - Added explicit CLI options
--d8-pathand--node-pathtobench_press runandbench_press validate, supporting custom binary overrides,D8_PATHandNODE_BINARY/NODE_EXECUTABLEenvironment variables, and SDK auto-probing for bundled D8 underbin/resources/dart2wasm/d8(Issue #19). - Fixed silent crashes on Node.js for JS and Wasm targets by replacing
stdout.writelnwithprint, generating self-invoking.run.mjsand.node.cjswrappers with unhandled rejection listeners, and forwarding CLI arguments viadartMainRunner(Issue #29, #30). - Added parameterized matrix group builder
BenchmarkGroup.matrix<T>(and convenienceBenchmark.matrix<T>) andBenchmarkMatrix<T>to benchmark competing implementations across parameterized inputs or datasets without repetitive boilerplate (Issue #23). - Added
mainBenchmarkMatrixCLI entrypoint and updatedmainBenchmarkSuiteto executeBenchmarkMatrixinstances seamlessly. - Added robust dispersion metrics to
BenchmarkMetrics: Median Absolute Deviation (madNs), normal-consistent robust CV (robustCv = (1.4826 * madNs) / medianNs), and Interquartile Range (iqrNs). - Added
isRobustStabletoBenchmarkMetricsand updatedisStableto incorporate robust dispersion, preventing transient bimodal GC sweeps from falsely failing steady-state stability for allocation-heavy workloads (Issue #20). - Added adaptive trial scaling via
--max-trialsCLI option andmaxTrialsinBenchmarkConfig/DefaultsConfig, allowingBenchmarkRunnerto dynamically collect additional measurement trials when initial variance exceeds threshold. - Enhanced
AdaptiveWarmupDetectorto distinguish systemic monotonic drift from transient bimodal outliers viacomputeRobustSemandhasSystemicDrift. - Added non-breaking
warmupComplete()lifecycle hook toBenchmarkandAsyncBenchmark.
0.2.0 #
-
Added multi-tier Cartesian comparison matrix support (Issue #5) via unified
bench_press.yamlconfiguration manifest,--config, and--dry-runinspection flag. -
Added N-dimensional
coordinates: Map<String, String>mapping toBenchmarkEntrytelemetry schema, replacing the single-axisgroupproperty. -
Added
MarkdownReporter.renderMatrixComparisonTableto render multidimensional matrix reports with grouped left-hand dimension columns and Fieller 95% ratio confidence intervals. -
Extracted mathematical and calibration constants (
Lanczos,Acklam, andBenchmarkCalibratorthresholds) with detailed doc comments. -
Removed legacy transitional
--compare-sdkoption in favor of unified Cartesian matrix configurations inbench_press.yaml. -
Added positional argument support (
<baseline> [current]) tobench_press diffalongside--baseline(-b) and--current(-c). -
Added
Blackhole.consumeStringandBlackhole.consumeObjectoverloads with@pragma('dart2js:never-inline')compiler barriers. -
Added
maxSemRelativeErroroption (default0.03) toBenchmarkConfigfor steady-state warmup convergence. -
Added
BenchmarkEntry.copyWithmethod and updatedBenchmarkEntry.keyto include optionalgroup($name:$target:$group) with deterministic name/target ordering inBenchmarkSuiteResult.deepMerge. -
Updated
BenchmarkSuiteResult.groupsto return group names in order of appearance rather than sorted alphabetically. -
Aligned default
targetBatchDurationto100msacross CLI runners andBenchmarkConfig. -
Unified default benchmark discovery directories across
runandvalidatecommands (benchmark,benchmarks,bench). -
Simplified benchmark discovery to convention-based matching (targeting files ending in
*_benchmark.dartor*_bench.dart) requiring standardvoid main()entrypoints, eliminating ad-hoc regex content parsing and dynamic wrapper script generation. -
Updated benchmark discovery to default strictly to
benchmark/(orbenchmarks/,bench/) and throw explicit errors (FormatExceptionfor non-Dart files,PathNotFoundExceptionfor nonexistent paths) rather than silently ignoring files or walking the entire repository root. -
Removed
BenchmarkFileKindenum and dynamic wrapper script generation. -
Removed deprecated
KbssdWarmupDetectoralias in favor ofAdaptiveWarmupDetector. -
Fixed steady-state warmup convergence math using Standard Error of the Mean (SEM) relative error (
<= 3%) and stationarity checks. -
Hardened
Blackhole.drain()compiler barrier against whole-program Dead-Store Elimination across AOT, Wasm, and JavaScript. -
Fixed
BenchmarkCalibratorto support sub-10µs operations without throwingCalibrationException, while throwingCalibrationExceptionby default when maximum probe batches produce zero elapsed ticks (elapsedUs == 0) unlessforceRun: true(--force-run) is specified (which warns and continues). -
Added
mode('sync'vs'async') property toBenchmarkResult(whichBenchmarkEntry.fromResultnow inherits for JSON telemetry). -
Implemented continuous Student's t-distribution quantile calculation (regularized incomplete beta for
1 < df < 2and Hill's Algorithm 396 fordf > 2) for accurate Fieller confidence intervals across all degrees of freedom. -
Fixed unhandled exception propagation in Isolate execution mode (
BenchmarkProcessRunner). -
Prevented floating-point overflow in geometric mean speedup reporting via log-sum calculation.
-
Updated
FiellerInterval.computeto returnisValid: false(withNaNbounds) when sample size is degenerate (N < 2).
0.1.0 #
- Initial release of
bench_press: A modern, statistically sound, compiler-aware multi-runtime benchmarking framework for Dart and Flutter. - Multi-runtime execution support across JIT, AOT (
dart compile exe), WasmGC (dart compile wasm), and JavaScript (dart compile js). Benchmark,AsyncBenchmark,BenchmarkVariant, andBenchmarkGroupharnesses with lifecycle hooks (setup,run,teardown).Blackholedead-code elimination (DCE) sink to safely consume benchmark results without compiler dead-code stripping.Throughputmetric tracking for byte rates (B/s,KB/s,MB/s,GB/s) and element rates (items/s,records/s,tokens/s).- Automated batch calibration and steady-state warmup convergence detection.
- Statistical summary metrics (Mean, Median, Min, Max, StdDev, CV, p95, p99, Ops/sec) with Fieller 95% confidence intervals for variant ratios.
- Markdown reporting with side-by-side variant comparisons and before/after baseline diffing.
bench_pressCLI withrun,validate,report, anddiffsubcommands.