dart_tensor_preprocessing 1.0.0
dart_tensor_preprocessing: ^1.0.0 copied to clipboard
High-performance tensor preprocessing library for Flutter/Dart. NumPy-like transforms pipeline for ONNX Runtime inference.
Changelog #
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
1.0.0 - 2026-09-16 #
Fixed #
-
SIMD binary kernels validate input/output lengths in release builds before modifying the output.
-
SimdOps.copy uses native typed-list copying to preserve overlapping views and rejects unequal lengths outside debug builds.
-
Mutation dtype dispatch rejects strided destinations rather than silently changing a temporary contiguous copy. Pair dispatch still accepts strided input.
-
BufferPool ignores duplicate returns of an already pooled buffer, preventing the same object from being lent to two callers simultaneously.
-
SIMD normalization falls back to direct division when its reciprocal overflows or underflows, including vector blocks and scalar tails.
-
Fused normalization divides directly by std, avoiding reciprocal overflow and incorrect NaN for a zero numerator with a subnormal standard deviation.
-
Normalize and fused resize/normalize defensively freeze mean/std lists so later caller mutations cannot bypass validation or change pipeline results.
-
Fused resize/normalize validates rank and channel count during shape inference.
-
RandomErasing rejects non-finite parameters and validates inferred rank; document package-specific sampling and uniform random fill.
-
PositionalEncoding rejects non-finite/non-positive bases and validates shape inference; clarify additive encoding versus rotary or learned embeddings.
-
All PadOp modes preserve exact integer values and validate inferred rank.
-
Symmetric reflection now repeats safely for kernels/padding larger than the input. GaussianBlur avoids sigma-square underflow, rejects non-finite sigma, validates inferred rank and preserves exact integer identity for kernel size 1.
-
RandomCrop preserves integer values without converting through double and rejects invalid rank/oversized crops consistently during shape inference.
-
Horizontal/vertical flips preserve exact integer storage instead of converting through double. Random flips reject non-finite probabilities, and every flip validates rank consistently during execution and shape inference.
-
Hue half-turns use an exact complementary-channel formula, preventing normalized integer channels from truncating 1 to 0 due to HSV roundoff. Color adjustment/jitter reject non-finite factors and validate shape inference.
-
RGB/grayscale/HSV output shape inference now rejects invalid ranks and channel counts consistently with execution.
-
Clip preserves integer values already inside the bounds without rounding them through double, including int64 values above 2^53.
-
Float32 ScaleOp subtracts offset before scaling, preventing severe cancellation from its previous expanded formula. Atan2 validates full shapes and snapshots overlapping in-place operands. Clip rejects NaN bounds explicitly.
-
Integer division and nonnegative integer powers no longer lose precision through double conversion. Integer division truncates toward zero and rejects zero divisors before mutation. Fractional operand behavior remains explicit double arithmetic followed by integer truncation and destination conversion.
-
Preserve exact integer add/subtract/multiply for integer tensor operands and integral signed-64-bit scalar operands, avoiding double conversion above 2^53.
-
Pow exponents 0.5/-0.5 use sqrt/reciprocal-sqrt semantics, including NaN for negative infinity, matching the pinned PyTorch reference.
-
Binary arithmetic rejects tensor operands with different shapes even when element counts match. In-place tensor arithmetic snapshots the other operand so overlapping views use original values throughout the calculation.
-
PermuteOp shape inference rejects duplicate/out-of-range axes, consistent with execution. PermuteOp/ReshapeOp copy their parameter lists so external mutation cannot invalidate a previously validated operation.
-
LayoutConvertOp now always performs its documented directional NCHW/NHWC permutation, agreeing with shape inference and round-trip conversion. Physical channels-last strides no longer cause conversion to be skipped. sliceFirst preserves existing strides instead of reinterpreting layout metadata.
-
Vector select/unbind now return aliased [1] tensors instead of failing on rank-zero construction. Unbind reuses select and preserves memory metadata. Narrow rejects empty, negative and overflowing ranges before constructing views.
-
Copy and freeze shape/stride metadata, validate storage spans and nonnegative strides, and reject invalid reshape dimensions. Squeeze/unsqueeze now agree with computed shapes for negative axes; single-element squeeze retains [1] instead of creating an unusable rank-zero view. See the migration notes.
-
Allocate eye/linspace/arange using the requested dtype; reject non-finite or empty sequences explicitly. Document double-sequence integer truncation.
-
Preserve exact integer sources in TypeCastOp, including values beyond 2^53; retain the documented legacy rounding, clamping and wrapping rules.
-
Bound contiguous storage kernels to the tensor's offset and element count, including in-place operations, indexing, concatenation and isolate transport.
-
Correct preset HWC/NHWC input order; preserve existing batches and avoid dividing floating-point image inputs by 255 twice.
-
Fix Tanh overflow and unseeded random calls reusing time-based seeds.
-
Improve exact GELU error-function accuracy for float64.
-
Preserve exact int64/uint64 values in clone and strided contiguous copies, and exact integer values through gather/slice/stack/concat/split/where/tile/roll/top-k.
-
Preserve axis reduction dtypes, promote integer sums to int64, and reject integer axis means. Select argmin/argmax without converting integers to double; propagate NaN in extrema and rank NaN consistently in top-k.
-
Accumulate repeated roll dimensions and validate gather index dtype/shape.
-
Allocate float64 random output correctly; keep uniform output below one after float32 rounding, and use standard Box-Muller math without truncated tails.
-
Match PyTorch Lp normalization's max(norm, eps) denominator; propagate NaN for the infinity norm, and reject invalid epsilon/order parameters.
-
Match PyTorch nearest coordinates, bicubic coefficient/border handling, area adaptive-average bins and torchvision shortest-edge size truncation.
-
Match torchvision center-crop rounding and zero padding for oversized crops.
-
Accept numeric masks independently of the selected tensor's dtype.
Added #
- Floating-point bilinear/bicubic antialias resize, used by image presets.
- Pinned PyTorch/torchvision CPU golden generator and independent full-value tests, including remote PNG fixtures with encoded/decoded checksums.
- Linux/Windows/macOS Dart checks and separate network/oracle CI jobs.
These changes alter incorrect 0.9.0 outputs. See the migration notes in README. The 1.0.0 release is pending the remaining compatibility audit and release gates.
0.9.0 - 2026-04-05 #
Added #
ArgMaxOp/ArgMinOp- Argmax/argmin index reduction operations (Issue #34, PR #62):argmax()/argmin()global methods returning flat index (int)argmaxAxis()/argminAxis()axis reduction returningTensorBufferwithDType.int64ArgMaxOp/ArgMinOpTransformOp subclasses with ONNX metadata (opset 13)- Supports
keepDimsparameter and negative axis indexing - Handles non-contiguous tensors via stride-based access
- Equivalent to
torch.argmax()/torch.argmin()in PyTorch and ONNXArgMax/ArgMinoperators
0.8.5 - 2026-04-03 #
Added #
TopKOp- Top-K selection operation along a specified axis (Issue #35):- Selects the k largest or smallest values and their corresponding indices
- Equivalent to
torch.topk()in PyTorch and ONNXTopKoperator (opset 1+) - Returns
(TensorBuffer values, TensorBuffer indices)record viaapplyTopK() - Pipeline-compatible
apply()returns values only - Parameters:
k(required),axis(default: -1),largest(default: true),sorted(default: true) - Extension method:
tensor.topk(k, axis:, largest:, sorted:) - Uses introselect algorithm (quickselect + insertion sort) for O(n) average-case selection
- Int64 index output, float32/float64 dtype-specialized value paths
- Supports negative axis indexing
Changed #
- Code formatting - Reformatted entire codebase (94 files) with Dart 3.x tall style formatter
0.8.4 - 2026-03-28 #
Added #
- CI Pipeline - GitHub Actions CI workflow with three jobs (PR #59):
- Static Analysis (
dart analyze --fatal-infos) on Dart stable - Format Check (
dart format --set-exit-if-changed) on Dart stable - Tests (
dart test) on Dart 3.9.0 and stable - Triggers on push and pull request to
main/masterbranches
- Static Analysis (
Fixed #
- SIMD NaN handling - Fixed
clipoperation insimd_ops.dartto use scalar loop that correctly preserves NaN values instead of SIMD path that silently converted them - Grayscale test tolerance - Adjusted test tolerance to account for ITU-R BT.601 coefficients summing to 0.9999 instead of 1.0
- Nearest resize test - Fixed expected values to match
align_corners-style coordinate mapping implementation - Code quality - Fixed unused imports, unused parameters,
prefer_const_declarations, andnon_constant_identifier_nameslint warnings across test files - Code formatting - Auto-formatted 31 files to comply with
dart formatstandards
0.8.3 - 2026-03-22 #
Added #
- Color Space Conversion Operations - Standalone ops for color space conversion (Issue #33):
RgbToGrayscaleOp- RGB to grayscale using ITU-R BT.601 weighted coefficients (0.2989R + 0.5870G + 0.1140B)- Output shape changes:
[3,H,W]→[1,H,W],[N,3,H,W]→[N,1,H,W] - Equivalent to
torchvision.transforms.Grayscale(num_output_channels=1)/tf.image.rgb_to_grayscale
- Output shape changes:
RgbToHsvOp- RGB to HSV color space conversion with H,S,V all in [0,1] range- Equivalent to
tf.image.rgb_to_hsv
- Equivalent to
HsvToRgbOp- HSV to RGB color space conversion- Equivalent to
tf.image.hsv_to_rgb
- Equivalent to
- All ops support 3D
[C,H,W]and 4D[N,C,H,W]tensors with dtype-specialized paths (Float32/Float64) - Reuses existing
rgbToHsv/hsvToRgbutility functions fromcolor_jitter_op.dart
0.8.2 - 2026-02-20 #
Added #
-
ErrorMessageshelper class - Static string formatting methods for consistent error messages:rankMismatch,rankRange,channelMismatch,shapeMismatch,axisBounds,dtypeMismatch,parameterRange,broadcastIncompatible
-
OpValidatorextensions - 3 new validation methods:validateImageTensor()- Validates 3D/4D image tensors with optional channel checkvalidateBroadcast()- Validates and computes broadcast output shape (NumPy rules)validateFloatDType()- Validates float32/float64 data types
-
SELUOp- Scaled Exponential Linear Unit activation (PyTorchF.selu, ONNXSelu)- Fixed constants: alpha=1.6732632423543772, scale=1.0507009873554805
- In-place support, dtype-specialized
-
GLUOp- Gated Linear Unit activation (PyTorchF.glu, ONNXSplit+Sigmoid+Mul)- Splits input along specified dimension and applies
a * sigmoid(b) - Output dimension halved along split axis
- Splits input along specified dimension and applies
-
LpNormalizeOp- Lp normalization along a specified dimension (PyTorchF.normalize, ONNXLpNormalization)- Factory constructors:
.l1(),.l2(),.linf() - Supports arbitrary p-norm values
- In-place support, dtype-specialized
- Factory constructors:
-
tensorWhere()/WhereOp- Element-wise conditional selection (PyTorchtorch.where, ONNXWhere)- Top-level function:
tensorWhere(condition, x, y) - TransformOp form:
WhereOp(condition:, y:)where input acts as x
- Top-level function:
-
MaskedFillOp- Fill tensor positions where mask is true (PyTorchTensor.masked_fill_)- Factory:
.attentionMask(mask)fills masked positions with negative infinity - In-place support
- Factory:
-
GatherOp- Gather elements along a dimension by index (PyTorchtorch.gather, ONNXGatherElements)- PyTorch-compatible indexing semantics
-
split()/chunk()- Tensor splitting utilities:split(tensor, splitSizes, {dim})- Split into specified sizeschunk(tensor, chunks, {dim})- Split into equal-sized chunks
-
TileOp- Tile/repeat tensor contents (ONNXTile, PyTorchTensor.repeat) -
RepeatOp- Repeat tensor (PyTorch.repeat()semantics, delegates to TileOp) -
RollOp- Circular shift along dimensions (PyTorchtorch.roll)- Supports per-dimension shifts and flat roll without dims
0.8.1 2026-02-20ÏÏ #
Added #
-
FloorOp,CeilOp,RoundOp- Element-wise rounding operations:FloorOp- Rounds down usingfloorToDouble()(PyTorchtorch.floor)CeilOp- Rounds up usingceilToDouble()(PyTorchtorch.ceil)RoundOp- Rounds to nearest usingroundToDouble()(PyTorchtorch.round)- Useful for coordinate calculations, index computation, and quantization preprocessing
-
RandomHorizontalFlipOp/RandomVerticalFlipOp- Random flip augmentation operations:HorizontalFlipOp- Deterministic left-to-right flipVerticalFlipOp- Deterministic top-to-bottom flipRandomHorizontalFlipOp- Probabilistic horizontal flip (default p=0.5)RandomVerticalFlipOp- Probabilistic vertical flip (default p=0.5)- Supports 3D
[C,H,W]and 4D[N,C,H,W]tensors with dtype-specialized paths - Optional seed parameter for reproducibility
-
RandomErasingOp- Random erasing (cutout) augmentation (PyTorchtorchvision.transforms.RandomErasing):- Configurable probability, scale range, and aspect ratio range
- Constant or random fill value for erased regions
- Supports 3D
[C,H,W]and 4D[N,C,H,W]tensors - In-place support, optional seeding for reproducibility
-
ColorJitterOp- Color jitter data augmentation with sub-operations:AdjustBrightnessOp- Additive brightness adjustment with clampingAdjustContrastOp- Per-channel mean-based contrast adjustmentAdjustSaturationOp- HSV-based saturation adjustmentAdjustHueOp- HSV-based hue rotation with wrappingColorJitterOp- Combined PyTorch-style random color jitter with seed support- RGB/HSV color space conversion utilities
- Supports 3D
[C,H,W]and 4D[N,C,H,W]tensors with 3-channel RGB validation
-
Trigonometric operations - Element-wise trig functions:
SinOp,CosOp,TanOp,AsinOp,AcosOp,AtanOp,Atan2Op- Follows existing
UnaryMathOppattern withOperationCapabilitiesmetadata - ONNX op types (opset 7+) and PyTorch equivalents
-
PositionalEncodingOp- Transformer positional encoding support
Fixed #
RoundOpdocstring - Corrected documentation: Dart uses half-away-from-zero rounding, not half-to-even (banker's rounding) like PyTorch
0.8.0 - 2026-02-15 #
Added #
-
CoordinateTransformModeenum - ONNX-compatible coordinate transformation modes forResizeOp:halfPixel- PyTorch default ((x + 0.5) * scale - 0.5)alignCorners- PyTorchalign_corners=Trueasymmetric- TensorFlow default (x * inSize / outSize)pytorchHalfPixel- Same as halfPixel but maps to 0 when outSize == 1- New
coordinateModeparameter onResizeOp(backward compatible with existingalignCornersbool)
-
OperationCapabilitiesexpanded - 5 new metadata fields for framework compatibility:supportsBroadcast- Whether the operation supports tensor broadcastingsupportedDTypes- Set of supported data types (default:{float32, float64})pytorchEquivalent- Equivalent PyTorch operation nameonnxOpType- Equivalent ONNX operator typeonnxOpsetVersion- Minimum ONNX opset version required
Changed #
- File decomposition - Large operation files split into focused modules:
activation_op.dart(1067 lines) →activation/subdirectory with 7 focused files:relu_ops.dart(ReLUOp, LeakyReLUOp)sigmoid_ops.dart(SigmoidOp, HardsigmoidOp, TanhOp)softmax_op.dart(SoftmaxOp)gelu_op.dart(GELUOp)swish_ops.dart(SiLUOp, SwishOp, HardswishOp)mish_op.dart(MishOp)elu_op.dart(ELUOp)
CenterCropOpextracted fromresize_op.darttocrop_op.dart- Barrel re-exports maintain backward compatibility for existing imports
Migration Notes #
ResizeOp: The newcoordinateModeparameter defaults tonull, preserving existing behavior viaalignCornersbool. No code changes needed for existing users.OperationCapabilities: All new fields have default values. Existingconst OperationCapabilities(...)calls remain valid.- File split:
activation_op.dartandresize_op.dartre-export all symbols. Existingimportstatements continue to work.
0.7.0 - 2026-02-02 #
Added #
-
New Activation Functions (PyTorch compatible):
GELUOp- Gaussian Error Linear Unit, standard in Transformers (BERT, GPT, ViT)- Supports exact computation and
tanhapproximation modes
- Supports exact computation and
SiLUOp(Swish) - Sigmoid Linear Unit, used in EfficientNet and YOLOv5SwishOp- Alias forSiLUOpHardsigmoidOp- Hardware-efficient sigmoid approximation for MobileNetV3HardswishOp- Hardware-efficient swish approximation for MobileNetV3MishOp- Self-regularizing activation used in YOLOv4+ELUOp- Exponential Linear Unit with configurable alpha
-
stack()Function - Stack tensors along a new dimension (torch.stack equivalent)- Supports arbitrary dimension insertion with negative indexing
- All input tensors must have identical shapes
- Dtype-specialized for Float32/Float64 performance
-
New Normalization Operations (PyTorch compatible):
InstanceNormOp- Instance normalization for style transfer and GANs- Normalizes per sample per channel (each spatial region independently)
- Supports 3D
[C,H,W]and 4D[N,C,H,W]tensors InstanceNormOp.fromStateDict()factory for loading PyTorch weights- Equivalent to
torch.nn.InstanceNorm2d
RMSNormOp- Root Mean Square normalization for modern LLMs- More efficient than LayerNorm (no mean subtraction)
- Used in LLaMA, Gemma, and other modern transformers
- Factory presets:
llama7B,llama13B,llama70B,gemma2B RMSNormOp.fromStateDict()factory for loading weights- Equivalent to
torch.nn.RMSNorm(PyTorch 2.4+)
Documentation #
- Updated PyTorch compatibility table in README.md
- Added new activation functions to Available Operations list
0.6.5 - 2026-02-02 #
Added #
-
OperationCapabilities metadata - All operations with
InPlaceTransformandRequiresContiguousmixins now overridecapabilitiesgetter:ReLUOp,LeakyReLUOp,SigmoidOp,TanhOp,SoftmaxOpUnaryMathOp(AbsOp, NegOp, SqrtOp, ExpOp, LogOp)ArithmeticOp(AddOp, SubOp, MulOp, DivOp),PowOpBatchNormOp,LayerNormOp,GroupNormOpNormalizeOp,ScaleOp,ClipOpResizeOp,CenterCropOp,RandomCropOp,GaussianBlurOp,PadOpTypeCastOp,ToTensorOp,ToImageOp
-
NaN/Infinity edge case tests - 23 new tests in
simd_ops_test.dart:- Float32/Float64 NaN handling for clip, abs, sqrt, normalize, relu operations
- Float32/Float64 Infinity handling for clip, abs, sqrt, normalize, relu operations
- Op-level NaN/Inf handling tests for ClipOp, AbsOp, SqrtOp, ReLUOp
Changed #
- Code consistency - Standardized
cloneForModification()usage across all in-place operations:BatchNormOp,LayerNormOp,ClipOp,ArithmeticOp,PowOpnow usecloneForModification()- Eliminates potential double-copy issues from manual contiguity checks
Performance #
PowOpdtype specialization - Added Float32/Float64 specialized loops for direct TypedList access
Documentation #
- Time/space complexity - Added Big-O complexity documentation to key operations:
ResizeOp- Complexity table for all interpolation modes (nearest, bilinear, bicubic, area, lanczos)NormalizeOp- O(n) time with SIMD accelerationBatchNormOp- O(n) time with pre-computed coefficientsLayerNormOp- O(n) time with Welford's algorithmGroupNormOp- O(n) time with per-group normalizationSoftmaxOp- O(n) time with 3-pass algorithmGaussianBlurOp- O(C×H×W×k) time using separable convolutionResizeNormalizeFusedOp- O(C×H_out×W_out) with no intermediate tensor
Tests #
- Total test count: 897 (23 new NaN/Inf edge case tests)
0.6.4 - 2026-01-28 #
Added #
ResizeNormalizeFusedOp- Fused resize + normalize operation that eliminates intermediate tensor allocation- Combines bilinear resize and per-channel normalization in a single pass
factory ResizeNormalizeFusedOp.imagenet(...)convenience constructor- Supports 3D
[C, H, W]and 4D[N, C, H, W]inputs - Cache-friendly 64x64 blocking for optimal L1 cache usage
Changed #
- Cache-friendly blocking for bilinear resize - Applied 64x64 blocking pattern to
_resizeBilinear()for both Float32-specialized and generic fallback paths - Cache-friendly blocking for area resize - Applied 64x64 blocking pattern to
_resizeArea()for both Float32-specialized and generic fallback paths ResizeNormalizeFusedOp.name- Now includesalignCornersparameter for better debugging visibility- Generic path style consistency - Pre-computes
oneMinusFy/oneMinusFxin_bilinearNormalizeGenericmatching Float32 path style
Tests #
- Added edge case tests for
ResizeNormalizeFusedOp: 1x1 input, same-size resize, alignCorners with dim=1, 25x upscale - Added validation tests: negative width, 5D input rejection,
alignCornersin name - Added path coverage tests: 4D+alignCorners, 4D+Float64 generic fallback, factory default alignCorners
- Added shape coverage tests: 4D non-contiguous, 4+ channel, batch=1 4D, computeOutputShape 2D behavior
0.6.3 - 2026-01-28 #
Added #
TensorBuffer.uninitialized()factory - Creates tensor buffer without zero-fill for cases where all elements will be immediately overwritten- Supports all DTypes and MemoryFormat options
- Semantically signals intent to overwrite, avoiding redundant initialization
Changed #
-
Uninitialized buffer usage - Operations that fully overwrite output now use
TensorBuffer.uninitialized()instead ofzeros():ResizeOp(3D/4D),CenterCropOp(3D/4D),concat(),SliceOp,RandomCropOp(3D/4D),GaussianBlurOp(3D/4D),PadOp(all modes)
-
BufferPool integration in GaussianBlurOp - Temporary
Float64Listbuffers now acquired fromBufferPooland properly released viatry/finallyto prevent leaks on exceptions
0.6.2 - 2026-01-20 #
Internal #
-
TensorBuffer Factory Separation - Moved factory methods to separate file:
tensor_buffer_factory.dartcontains: zeros, ones, full, random, randn, eye, linspace, arange, fromFloat32List, fromFloat64List, fromUint8List- Reduces
tensor_buffer.dartfrom ~840 lines to ~530 lines - No API changes
-
OpValidator - Added centralized operation validation (
validation_utils.dart):OpValidator.validateRank()- Validates tensor rank rangeOpValidator.validateAxis()- Validates and normalizes axis (supports negative indexing)OpValidator.validateChannels()- Validates channel countOpValidator.validatePositiveDimension()- Validates positive dimensionOpValidator.validateListLength()- Validates list length
-
OperationCapabilities - Added operation metadata (
transform_op.dart):supportsInPlace- Whether op can modify tensor in placerequiresContiguous- Whether op requires contiguous memorypreservesShape- Whether op preserves input shapemodifiesDType- Whether op may change data type- Default
capabilitiesgetter onTransformOp
0.6.1 - 2026-01-17 #
Added #
-
Float64 SIMD Operations - Vectorized operations for Float64 tensors (
simd_ops.dart):SimdOps.clipF64()- Clips values using Float64x2.clamp()SimdOps.absF64()- Absolute value using Float64x2.abs()SimdOps.sqrtF64()- Square root using Float64x2.sqrt()SimdOps.normalizeF64()- Mean/std normalization with SIMD- Uses Float64x2List.view() for aligned data (16-byte alignment)
- Scalar fallback for unaligned data to avoid object creation overhead
- ~2.5x speedup for aligned Float64 data vs scalar
-
SIMD Microbenchmark - Performance verification for SIMD operations (
benchmark/simd_microbenchmark.dart):- Direct SimdOps performance measurement (clip, abs, sqrt, normalize)
- Aligned vs unaligned data comparison (~4.4x performance difference)
- Float32 SIMD vs Float64 SIMD comparison
- Edge case testing for non-multiple-of-4 lengths
-
SIMD Tests - 64 tests in
simd_ops_test.dart:- Float32 SIMD tests with alignment edge cases
- Float64 SIMD tests (clipF64, absF64, sqrtF64, normalizeF64)
- Op integration tests for both Float32 and Float64
Changed #
- SimdOps.abs() and SimdOps.sqrt() - Now applied to
AbsOpandSqrtOpfor Float32 tensors - SimdOps.clip() - Now used in
ClipOpfor Float32 tensors - SimdOps.normalize() - Now used in
NormalizeOpfor Float32 tensors (per-channel) - NegOp - Now uses
SimdOps.multiplyScalar(-1)for Float32 tensors - ClipOp - Now uses
SimdOps.clipF64()for Float64 tensors - AbsOp - Now uses
SimdOps.absF64()for Float64 tensors - SqrtOp - Now uses
SimdOps.sqrtF64()for Float64 tensors - NormalizeOp - Now uses
SimdOps.normalizeF64()for Float64 tensors (3D and 4D)
Performance #
- Float32 SIMD (aligned): ~6.2 GE/s
- Float64 SIMD (aligned): ~3.3 GE/s (53% of Float32, expected due to Float64x2 vs Float32x4)
- Unaligned fallback: ~1.3-1.5 GE/s
Internal #
- Integrated SIMD microbenchmark into
benchmark/run_all.dart
0.6.0 - 2026-01-16 #
Added #
-
Multi-axis Reductions - Reduce along multiple axes at once (
tensor_buffer_reduce.dart):sumAxes(List<int> axes, {bool keepDims})- Sum along multiple axesmeanAxes(List<int> axes, {bool keepDims})- Mean along multiple axesminAxes(List<int> axes, {bool keepDims})- Min along multiple axesmaxAxes(List<int> axes, {bool keepDims})- Max along multiple axes- Supports negative axis indexing
- Validates duplicate axes
-
GroupNormOp - Group normalization for modern CNNs (
group_norm_op.dart):- Full PyTorch-compatible
torch.nn.GroupNormimplementation - Normalizes across groups of channels (used in U-Net, modern CNNs with small batch sizes)
- Supports 3D
[C,H,W]and 4D[N,C,H,W]tensors GroupNormOp.withAffine()factory for PyTorch-style initializationGroupNormOp.fromStateDict()factory for loading PyTorch weights- Welford's algorithm for numerically stable mean/variance computation
- Dtype-specialized loops for Float32/Float64
- In-place support via
applyInPlace()
- Full PyTorch-compatible
-
SIMD Operations - Vectorized tensor operations (
simd_ops.dart):- Uses Float32x4 SIMD instructions for 2-4x speedup on Float32 tensors
SimdOps.multiplyScalar(),SimdOps.addScalar(),SimdOps.subtractScalar()- Scalar operationsSimdOps.add(),SimdOps.subtract(),SimdOps.multiply(),SimdOps.divide()- Element-wise binary operationsSimdOps.relu(),SimdOps.leakyRelu()- Activation functionsSimdOps.normalize()- Mean/std normalizationSimdOps.copy(),SimdOps.fill(),SimdOps.sum(),SimdOps.clip()- Handles both aligned and unaligned memory
-
Interpolation Modes - Additional resize algorithms (
resize_op.dart):InterpolationMode.area- Weighted area averaging for high-quality downsampling with anti-aliasing (OpenCV INTER_AREA equivalent)InterpolationMode.lanczos- Lanczos3 (6x6 kernel) for high-quality resize with sinc-based interpolation
Changed #
- BREAKING: Reduction operations moved to extension (
TensorBufferReduce)sum(),mean(),min(),max()- Full tensor reductionssumAxis(),meanAxis(),minAxis(),maxAxis()- Single-axis reductionstoList()- Data extraction- Existing code using these methods will work unchanged, but users importing only
tensor_buffer.dartmust now also importtensor_buffer_reduce.dartor the main library
Performance #
- SIMD-accelerated operations:
ScaleOp,ReLUOp,LeakyReLUOpnow use SIMD for Float32 tensors - SIMD-accelerated ArithmeticOp:
AddOp,SubOp,MulOp,DivOpnow use SIMD for Float32 tensors (both scalar and tensor modes) - Cache-friendly bicubic resize: 64x64 block processing for better L1 cache utilization on large tensors
Internal #
- Extracted reduction operations from
tensor_buffer.dart(1170 → 740 lines) totensor_buffer_reduce.dart - Added 14 new tests for multi-axis reductions
- Added 61 new tests for SIMD operations, GroupNormOp, and resize modes
PyTorch Compatibility #
| Operation | PyTorch Equivalent |
|---|---|
GroupNormOp |
torch.nn.GroupNorm |
0.5.1 - 2026-01-13 #
Added #
-
BufferPool - Memory pooling API for buffer reuse (
buffer_pool.dart):- Singleton
BufferPool.instancefor global buffer reuse - Power-of-2 size bucketing for efficient allocation
- Per-dtype buffer pools (Float32, Float64, Int32, Uint8, etc.)
acquire(minSize, dtype)andrelease(buffer)methodsacquireFloat32(),acquireFloat64(), etc. convenience extensions- Max buffers per bucket limit (8) to prevent unbounded memory growth
pooledCountandpooledBytesfor monitoring
- Singleton
-
TypedData Views - Zero-copy tensor view utilities (
typed_data_views.dart):TypedDataViews.float32SublistView()- Zero-copy Float32List slicingTypedDataViews.float64SublistView()- Zero-copy Float64List slicingTypedDataViews.viewAs()- Create typed view from ByteBuffer at offsetTensorViewExtensionon TensorBuffer:sliceFirst(start, end)- Zero-copy slice along first dimensionisViewable- Check if tensor can be used as a viewtoChannelsLast()- NCHW to NHWC without copyingtoChannelsFirst()- NHWC to NCHW without copyingflatten()- 1D view of contiguous tensorunbind(dim)- Split tensor into views along dimensionselect(dim, index)- Select single index with reduced ranknarrow(dim, start, length)- Narrow dimension without copying
-
Utility Libraries (
lib/src/utils/):dtype_dispatcher.dart- DTypeDispatcher for dtype-specialized dispatchtensor_indexing.dart- TensorIndexer for index calculations (index2D, index3D, index4D, linearToCoords, coordsToLinear, computeStrides)
-
TensorBuffer/TensorStorage Factory Methods:
TensorBuffer.fromFloat64List()- Create tensor from Float64ListTensorStorage.fromFloat64List()- Create storage from Float64List
Changed #
-
SoftmaxOp Optimization: Now preserves input dtype (Float32/Float64) instead of always using Float64. Added dtype-specialized implementations for better performance.
-
Double-copy elimination: Operations now use
cloneForModification()pattern (input.isContiguous ? input.clone() : input.contiguous()) to avoid unnecessary copies:ReLUOp,LeakyReLUOp,SigmoidOp,TanhOp,SoftmaxOpAbsOp,NegOp,SqrtOp,ExpOp,LogOp(UnaryMathOp)NormalizeOp,ScaleOp
Internal #
- Added
cloneForModification()helper toRequiresContiguousmixin intransform_op.dart - Integrated
DTypeDispatcherinto activation ops (ReLUOp,LeakyReLUOp,SigmoidOp,TanhOp) for dtype-specialized loops - Integrated
DTypeDispatcherintoScaleOpfor consistent dtype handling - Replaced stride computation with
TensorIndexer.computeStrides()inSoftmaxOp(removed 3x code duplication)
0.5.0 - 2026-01-10 #
Added #
-
BatchNormOp - Batch normalization for CNN inference (
batch_norm_op.dart):- Full PyTorch-compatible
torch.nn.BatchNorm2dimplementation - Pre-computed scale/shift coefficients for efficient inference:
y = x * scale + shift - Supports 3D
[C,H,W]and 4D[N,C,H,W]tensors BatchNormOp.fromStateDict()factory for loading PyTorch weights- Dtype-specialized loops for Float32/Float64
- In-place support via
applyInPlace()
- Full PyTorch-compatible
-
LayerNormOp - Layer normalization for Transformer inference (
layer_norm_op.dart):- Full PyTorch-compatible
torch.nn.LayerNormimplementation - Normalizes over last N dimensions (e.g.,
[768]for BERT) - Welford's algorithm for numerically stable mean/variance computation
LayerNormOp.bert()andLayerNormOp.bertLarge()factory presetsLayerNormOp.fromStateDict()factory for loading PyTorch weights- Dtype-specialized loops for Float32/Float64
- In-place support via
applyInPlace()
- Full PyTorch-compatible
PyTorch Compatibility #
| Operation | PyTorch Equivalent |
|---|---|
BatchNormOp |
torch.nn.BatchNorm2d (inference) |
LayerNormOp |
torch.nn.LayerNorm |
0.4.1 - 2026-01-09 #
Performance Optimizations #
-
Dtype-specialized loops: Hot paths in transform operations now use dtype-specific code paths with direct
Float32List/Float64Listaccess, avoiding per-element switch overhead:NormalizeOp._normalize3D(),NormalizeOp._normalize4D()ScaleOp._scale()ClipOp._clip()GaussianBlurOp._applySeparableBlur()ResizeOp._resizeNearest(),_resizeBilinear(),_resizeBicubic()CenterCropOp._crop3D(),_crop4D()concat()with optimized axis=0 bulk copy
-
Clone-Before-Modify optimization:
ClipOp.apply()now avoids double copy by checkingisContiguousbefore deciding whether toclone()orcontiguous() -
Isolate threshold:
TensorPipeline.runAsync()now accepts optionalisolateThresholdparameter (default: 100,000 elements). Small tensors skip isolate overhead and run synchronously -
Buffer reuse:
GaussianBlurOpnow pre-allocates and reuses temp buffer across channels, reducing allocations -
Concat linear copy:
concat()now uses pre-computed strides for linear index calculation instead of recursive index computation. Axis=0 concatenation of contiguous tensors uses bulksetRange()copy -
Loop unrolling:
ResizeOp._resizeBicubic()unrolls 4x4 kernel with pre-computed weights and indices
0.4.0 - 2026-01-09 #
Added #
- Arithmetic Operations (
arithmetic_op.dart):AddOp- Element-wise addition (scalar or tensor)SubOp- Element-wise subtraction (scalar or tensor)MulOp- Element-wise multiplication (scalar or tensor)DivOp- Element-wise division (scalar or tensor)PowOp- Element-wise power operation
- Math Operations (
math_op.dart):AbsOp- Element-wise absolute valueNegOp- Element-wise negationSqrtOp- Element-wise square rootExpOp- Element-wise exponential (e^x)LogOp- Element-wise natural logarithm
- Activation Functions (
activation_op.dart):ReLUOp- Rectified Linear UnitLeakyReLUOp- Leaky ReLU with configurable negative slopeSigmoidOp- Sigmoid activationTanhOp- Hyperbolic tangent activationSoftmaxOp- Softmax along specified axis
- TensorBuffer Factory Methods:
TensorBuffer.full()- Create tensor filled with specified valueTensorBuffer.random()- Create tensor with uniform random values [0, 1)TensorBuffer.randn()- Create tensor with standard normal distributionTensorBuffer.eye()- Create identity matrix (supports rectangular)TensorBuffer.linspace()- Create tensor with evenly spaced valuesTensorBuffer.arange()- Create tensor with sequence values
- Utility Libraries (
lib/src/utils/):index_utils.dart- Index manipulation utilities (reflectIndex, replicateIndex, circularIndex)validation_utils.dart- Common tensor validation patterns
Changed #
- Exception Consistency:
TensorStorage._checkBounds()now throwsIndexOutOfBoundsExceptioninstead ofRangeErrorfor consistent exception handling across the library
Internal #
- Extracted duplicate
_reflectIndexcode frompad_op.dartandaugmentation_op.dartinto shared utility - Added
TensorValidationextension withrequireRank3Or4(),requireExactRank(),requireMinRank()methods
0.3.1 - 2026-01-08 #
Added #
- Performance benchmark suite (
benchmark/directory):tensor_creation_benchmark.dart- Tensor creation performancetensor_ops_benchmark.dart- Zero-copy and copy operationspipeline_benchmark.dart- Pipeline sync/async comparisonmemory_benchmark.dart- Memory usage measurementrun_all.dart- Unified benchmark runnerutils/benchmark_utils.dart- Benchmark utilities
Fixed #
- Removed unused variables in benchmark files
- Fixed lint issues in benchmark files
0.3.0 - 2026-01-08 #
Added #
ClipOp- Element-wise value clamping with factory presets (unit, symmetric, uint8)PadOp- Padding with multiple modes (constant, reflect, replicate, circular)SliceOp- Python-like tensor slicing with support for negative indices and stepsRandomCropOp- Random cropping for data augmentation with deterministic seed supportGaussianBlurOp- Gaussian blur using separable convolution with factory presetsconcat()- Utility function for tensor concatenation along specified axis
Fixed #
concat()axis-based copy logic now correctly handles multi-axis concatenation
Changed #
- BREAKING: Unified exception handling across the library
- All exceptions now extend
TensorExceptionsealed class ArgumentError→ShapeMismatchException,InvalidParameterExceptionRangeError→IndexOutOfBoundsException
- All exceptions now extend
0.2.0 - 2026-01-04 #
Added #
IndexOutOfBoundsException- Thrown when an index or axis is out of valid rangeDTypeMismatchException- Thrown when tensor data types do not match
Changed #
- BREAKING: Unified exception handling across the library
- All exceptions now extend
TensorExceptionsealed class ArgumentError→ShapeMismatchException,InvalidParameterExceptionRangeError→IndexOutOfBoundsExceptionStateError→NonContiguousException,DTypeMismatchException
- All exceptions now extend
- Shape validation now happens before buffer creation in
zeros()andones()
Migration Guide #
If you were catching standard Dart exceptions, update your code:
| Before | After |
|---|---|
on RangeError |
on IndexOutOfBoundsException |
on ArgumentError |
on ShapeMismatchException or on InvalidParameterException |
on StateError |
on NonContiguousException or on DTypeMismatchException |
0.1.4 - 2026-01-04 #
Added #
- Reduction operations for
TensorBuffer:sum()- Returns the sum of all elementsmean()- Returns the arithmetic mean of all elementsmin()- Returns the minimum valuemax()- Returns the maximum value
- Axis-wise reduction operations:
sumAxis(int axis, {bool keepDims})- Sum along a specific axismeanAxis(int axis, {bool keepDims})- Mean along a specific axisminAxis(int axis, {bool keepDims})- Min along a specific axismaxAxis(int axis, {bool keepDims})- Max along a specific axis
- Support for negative axis indexing in axis-wise operations
- Comprehensive test coverage for all reduction operations (49 tests)
0.1.3 - 2026-01-03 #
0.1.1 - 2025-12-27 #
Added #
- Comprehensive dartdoc comments for all public API elements
- Library-level documentation with usage examples
0.1.0 - 2025-12-27 #
Added #
-
Core tensor operations
TensorBufferwith shape, strides, and view/storage separationTensorStoragefor immutable typed data wrapperDTypeenum with ONNX-compatible data types
-
Transform operations
ResizeOpwith nearest, bilinear, bicubic interpolationResizeShortestOpfor aspect-ratio preserving resizeCenterCropOpfor center croppingNormalizeOpwith ImageNet, CIFAR-10, symmetric presetsScaleOpfor value scalingPermuteOpfor axis reorderingToTensorOpfor HWC uint8 to CHW float32 conversionToImageOpfor CHW float32 to HWC uint8 conversionUnsqueezeOp,SqueezeOp,ReshapeOp,FlattenOpfor shape manipulationTypeCastOpfor dtype conversion
-
Pipeline system
TensorPipelinefor chaining operationsPipelinePresetswith ImageNet, ResNet, YOLO, CLIP, ViT, MobileNet presets- Async execution via
Isolate.run
-
Zero-copy operations
transpose()via stride manipulationsqueeze(),unsqueeze()as shape-only changes