monowave 0.4.0
monowave: ^0.4.0 copied to clipboard
Headless audio for Flutter: microphone capture, waveform peaks, and non-destructive editing, from one C core across all six Flutter targets.
Changelog #
This file records all notable changes to monowave. The file uses the format of Keep a Changelog. The project uses semantic versioning.
0.3.0 is the first version on pub.dev. 0.1.0 through 0.2.0 were development milestones. This file keeps them, because what changed in them is still the history of this API.
0.4.0 #
monowave can play what it edits. The same C loop that the exporter runs also renders an edit. Therefore what you hear before you commit to a trim is what the export writes afterwards, byte for byte, on all six targets.
Breaking change for a host that implements MonowavePlatform #
Version 0.4.0 adds four members to the seam: renderPcm, renderPcmBytes,
openPlayback and openPlaybackBytes. The change does not affect code that
calls monowave. It also does not affect a host that uses
FakeMonowavePlatform, because the fake implements the four members. A host
that implements the interface itself must add the four methods.
Added: monowave can render an edit without a file #
MonowavePlatform.renderPcm returns a document as 16-bit PCM. The samples are
byte-identical to what exportWav writes for the same document. This equality
is deliberate rather than accidental. You can hear an edit before you commit to
it. Such a preview is useful only when what you hear is what you get.
The equality is structural. No test holds it in place. wf_export_wav is now a
file sink over a new wf_render, so the exporter and the renderer run one
loop. There is no second implementation that can drift.
The corpus behind that claim is the set of shapes that break a naive extraction:
- a zero-length region
- a negative-length region
- fades that are longer than the region that holds them
- a gain of exactly 1.0
- a gain that clamps
- a region that goes past the end of the source
- a region length that neither block size divides
The renderer reads in 1000-frame blocks, and the exporter reads in 4096-frame blocks. The output must still match. It can match only because the fade envelope depends on the position inside its region, not on the position of a block boundary.
wf_envelopeis now an exported symbol rather than astaticone. Therefore the exporter, the renderer and a future web implementation share one symbol for the curve. It also has the property tests that it never had. The tests assert these properties:- The envelope is exactly 0 and exactly 1 at the endpoints.
- The envelope is linear between the endpoints.
- Two fades that overlap multiply together.
- The multiplier is never negative.
FakeMonowavePlatform.renderPcmrecords every request and answers with silence of the right length, so a host can assert on the request without synthetic audio.renderPcmis not available on web, because web has no filesystem to read a source from.
This is M6 of the playback work. PlaybackSession and an audio device arrive
in M7 and M8. ROADMAP.md has more information.
Added: monowave can feed a document to an audio device #
wf_playback puts the M6 renderer behind a lock-free SPSC ring and a feeder
thread, and hands the result to a miniaudio output device. It is the capture
pipeline in the opposite direction, and it obeys the same rule. The audio
callback copies from a ring and does nothing else. The callback does not
allocate memory, does not take a lock, and does not call into Dart.
The feeder is a C thread rather than a timer on the Dart side. Capture tolerates a late drain, because its ring holds seconds and a stalled consumer loses nothing. Playback does not tolerate a late drain. An empty ring is silence in the speaker. With a timer on the Dart side, one garbage collection makes an audible dropout.
The session counts underruns rather than hides them, in the same way as
CaptureSession.dropped. Silence past the end of the render is the end of the
render rather than a fault. The counter knows this difference.
wf_playback_pull is public for the same reason as wf_capture_feed. It is
the audio-thread entry point, so a test can drive the realtime path on every
platform with no sound card. The test drives 30 seconds of an edited document
through the ring. The test then asserts that the result is byte-identical to
what M6 renders offline, with zero underruns.
This is M7. PlaybackSession, a seek and a position clock arrive in M8.
ROADMAP.md has more information.
src/wf_playback.ccarries the first per-platform code in the package. The code is a thread, a join and a sleep behind#if defined(_WIN32). miniaudio keeps its own thread abstraction private (ma_thread_createisstatic). A second vendored threading library for three functions was the worse trade.- Playback is not in the WASM build. Playback reads through the path-based decoders that this build compiles out, and a single-threaded module has nowhere to put a feeder.
Added: PlaybackSession, with a seek and a playhead #
MonowavePlatform.openPlayback returns a PlaybackSession: play, pause,
seek, position, duration, isPlaying, isFinished and underruns. It
is the mirror of CaptureSession, and what it plays is byte-identical to what
exportWav writes for the same document.
The playhead is the device, not a timer. position comes from the frames
that the device actually consumed. A timer drifts against the audio clock
immediately, and the playhead then sits where the sound is not. The
DemoPlayer in the example does this, correctly for a fake and wrongly for
anything real. A test asserts that the playhead does not move while the device
consumes nothing.
A seek is a handshake. A feeder fills the ring and a device drains it.
Neither of them can flush the ring, because neither of them can block.
Therefore wf_playback_seek blocks its own caller instead. It waits for the
consumer to leave the audio callback and for the feeder to park. It resets the
ring only after both of these events.
The device receives a few milliseconds of silence, and that is the right sound for a seek. The accuracy of a seek is the accuracy of the decoder. A seek is sample-exact on WAV. On MP3 a seek is accurate to approximately 26 ms. The seek goes to a frame boundary and then decodes forward.
PlaybackSession deliberately does not implement MonoPlaybackController.
That interface lives in monokit, and a headless package must not depend on the
design system. Therefore the adapter stays the few lines that a host writes.
package:monowave/testing.dartnow hasFakePlaybackSession. A test asserts that the fake and the real session agree on the shared contract. The contract includes a seek past the end, which both clamp, and counters, which both answer after disposal. Therefore a host cannot build a scrubber against the fake and then meet a different contract on a device.FfiPlaybackSessioncarries aNativeFinalizerfrom the start. A discarded session releases the output device and the feeder thread. It does not wait for the process to exit.
This is M8. Live document updates are M9, and web playback is M10. ROADMAP.md has more information.
Added: monowave can swap a document while it plays #
PlaybackSession.setDocument replaces the region list of a session while the
session runs. A user can drag a trim handle or change a gain, and hear the
result while playback continues. This is the feature that the whole playback
path exists for.
The playhead keeps its position in the output timeline. Therefore the sound
continues from that position and does not jump. If a change makes the document
shorter than the playhead position, the playhead clamps to the new end. A host
that wants different behavior calls seek immediately after the swap.
One call rather than two. The roadmap did not decide whether a gain change and a trim must be separate entry points. The theory was that a gain change is cheaper, but it is not. Both operations erase the frames in the ring, because the renderer made those frames from the old document. The ring holds approximately one second. A second entry point gives a host nothing except one more decision to get wrong.
Mechanically, a swap is a seek with a new region list. Therefore it reuses the seek handshake from M8 unchanged.
FakePlaybackSession.setDocumentrecords every swap and clamps in the same way, and a test asserts that the two agree.- After a rejected swap, the old document continues to play. C copies the new list before it releases the old list, so a failed allocation changes nothing.
This is M9. Web playback is M10. ROADMAP.md has more information.
Added: web renders through the same C loop, byte for byte #
MonowavePlatform.renderPcmBytes renders a document from bytes that are
already in memory. It is the render path on every target. It is the only render
path on web, because web has no filesystem.
The roadmap expected this to cost the equality guarantee. It did not. The
plan was to decode through the browser and to apply the envelope in an
AudioWorklet. That plan leaves MP3 approximate on web, because the decoder of
Chrome is not dr_mp3. Version 0.4.0 withdraws that concession.
DR_WAV_NO_STDIO and its siblings remove only the file entry points.
Therefore the WASM build already carried all three decoders with their memory
APIs, and wf_decode_memory always used them. Web now runs the same render
loop as the exporter, over the same decoders.
A rendered document is byte-identical on all six targets, for every format.
tool/verify_wasm.mjs asserts it. It renders a document through the WASM
module and compares every sample against a native render. A change of the int16
conversion from truncation to rounding fails that check at sample 1.
- Only
wf_source_openandwf_export_wavsit behind the stdio guard now. All targets share the source layer, the region walk, the envelope and the render loop. wf_region_strideis now an exported symbol. Therefore a binding that builds the array by hand (the web binding writes into the WASM heap) asks for the stride rather than guesses it.- monowave copies the bytes into the render. dr_libs reference the buffer of the caller rather than copy it, and a render outlives the call that made it.
An audio device is still missing on web. openPlayback throws there. The
difficult half is complete, because web can produce exactly the right samples.
What remains is a WebAudio graph to play them, and that graph needs no C.
This is M10. ROADMAP.md has more information.
Added: playback on web #
MonowavePlatform.openPlaybackBytes returns a PlaybackSession on all six
targets. Web has no filesystem, so bytes are the way in there. On native it is
the same engine as openPlayback. The engine reads from a copy rather than
streams a file.
What plays is byte-identical to what exportWav writes on every target,
because every target renders through the same C loop. Only the device is
different. Native targets use miniaudio, and web uses a WebAudio graph.
The web graph is deliberately plain. The whole render goes into an
AudioBuffer at the start, and an AudioBufferSourceNode plays it. There is
no ring and no feeder, because the browser owns the audio thread and there is
nothing to race. This design is correct for a preview of an edit and wrong for
an audiobook. A long document costs its whole length in memory.
underruns is always zero on web, and that is a fact rather than a stub. The
render stays in memory, so a feeder cannot lose a race that it is not in.
The playhead is still the audio clock. AudioContext.currentTime advances with
the hardware. The native session obeys the same rule, and counts the frames
that the device consumed.
Browsers refuse to start an AudioContext without a user gesture. A
browser reports this as a context that stays in suspended rather than as an
error. play() detects that state and throws PlaybackUnavailable. The error
message tells the host to call play() from a tap.
wf_playback_create_memoryis the native half, so a host that holds audio in memory does not have to write a temporary file to play it.openPlaybackwith a path still throws on web.- The M7 to M9 transport does not port, by design. A browser has its own scheduler, and only the samples cross.
Changed #
- ABI 13 → 14. The change is additive:
wf_playback_create_memory. - ABI 12 → 13. The change is additive:
wf_render_open_memoryandwf_region_stride. - ABI 11 → 12. The change is additive:
wf_render_set_regionsandwf_playback_set_regions. - ABI 10 → 11. The change is additive:
wf_render_seekandwf_playback_seek. - ABI 9 → 10. The change is additive: the
wf_playback_*surface. No existing signature changed. - ABI 8 → 9. The change is additive:
wf_envelope,wf_render_open,wf_render_close,wf_render_read,wf_render_sample_rate,wf_render_channelsandwf_render_length_frames. No existing signature changed, and the determinism digests do not move on either binding.
0.3.1 #
A capture session that kept the microphone open after it was dropped, an abandoned recording left corrupt on disk, and an RMS series that reached five of the six targets.
Fixed: a dropped capture session kept the microphone open #
FfiCaptureSession had no finalizer, so wf_capture_destroy ran only from an
explicit dispose(). A consumer that dropped a session without one leaked the
wf_capture struct, both lock-free rings, the preallocated take history and the
two drain buffers - and, worse, left the platform input device open, so the
microphone stayed live for the rest of the process. The decoded-peaks path had
had a NativeFinalizer since it was written; capture never got one.
It does now, over wf_capture_destroy, which stops the device before it frees
anything. Two things had to change for it to be able to do its job:
- The drain buffers moved into C. They were
calloced from Dart and freed indispose(), which is exactly the memory a finalizer cannot reach: aNativeFinalizeroverwf_capture_destroyfrees what the C struct owns and nothing else.wf_capture_scratchandwf_capture_pcm_scratchnow hand out buffers allocated with the session, so everything it owns hangs off one pointer and one call reclaims all of it. - The drain timer holds the session weakly. A pending
Timer.periodicis a GC root, and its callback capturedthis- so a session dropped while recording stayed reachable forever and would never have been collected at all. That is the case where the leak matters most, because the device is still open. The timer now cancels itself on the first tick after the session goes away.
dispose() remains the contract, and the documentation still says to call it.
A finalizer runs whenever the collector gets to the object, which may be long
after the recording ended and is not guaranteed before the process exits.
Nothing but dispose() releases the microphone promptly.
Also fixed alongside it: an unwritable CaptureConfig.recordTo path leaked the
session outright, because it threw between wf_capture_create and the
constructor - with no Dart object in existence yet for a finalizer to attach to.
Fixed: disposing a session left a corrupt recording and a freed pointer #
Two loose ends around CaptureSession.dispose, both found while adding the
finalizer above, and both behaviour changes in their own right.
It never closed the recording file. stop() rewrites the WAV header with
the real sizes and closes the file; dispose() did neither. Opening with
CaptureConfig.recordTo and disposing without stopping - which is what
cancelling a recording looks like - leaked the open handle and left a WAV whose
header still claimed the audio after it was zero bytes long, so every player
opened it and showed nothing. dispose() now closes it the same way stop()
does, and an abandoned take is playable rather than corrupt. It deliberately
does not drain first: finishing a take is still stop()'s job, and whatever the
audio thread published since the last pass is lost. The close is idempotent, so
the ordinary stop() then dispose() sequence is unaffected.
produced, dropped, pcmDropped and truncated read freed memory after
it. All four read straight out of the C struct, which dispose() frees. They
now answer from a tally frozen at disposal. Frozen rather than throwing, unlike
start or feedSynthetic: a counter is a query with a correct answer after
disposal, and FakeCaptureSession already answered it - so a host whose
visualizer read one while tearing down would have passed its widget tests and
then read freed memory against a real microphone. isRecording and isPaused
are now false after disposal too, on both the real session and the fake.
Fixed: peaks.rms(level) was null on web, and only on web #
RMS landed in 0.3.0 and reached five of the six targets. wf_peaks_rms was
never added to -sEXPORTED_FUNCTIONS in tool/build_wasm.sh, the _Core
extension type in lib/src/platform/wasm_platform.dart never declared it, and
_copyOut built its WaveformPeaks without an rms: argument - so on web
every level answered null while the same C source, over dart:ffi, returned
the series. A two-layer waveform lost its core on the one target that could not
be checked against a digest, and the package's claim to be byte-identical
everywhere was, for that release, not true.
All three are fixed, and the artifact did not have to change: WF_EXPORT marks
the symbol visibility("default"), so wf_peaks_rms was in the committed
assets/monowave.wasm export table the whole time. Nothing was missing from the
binary. What was missing was anything that looked.
The determinism check now covers it, because it did not. The FNV-1a digest
in test/fixtures.dart hashed the min/max levels alone, so the one check whose
job is to prove the two bindings agree was blind to an entire series - it stayed
green through the whole of 0.3.0. It now takes a WaveformPeaks and folds in
both series at every level, and it throws rather than skipping when a level has
no RMS, since a binding that never read the series is exactly what it exists to
catch. tool/verify_wasm.mjs folds the same two series in the same order. The
pinned digests all move as a result; regenerate with
dart run tool/print_digests.dart.
Three checks now stand between this and a repeat, one per way it broke.
tool/verify_wasm.mjs asserts every function _Core declares is present in the
artifact, named in -sEXPORTED_FUNCTIONS, and actually called by the binding. A
member the extension type does not declare is a compile error, which is what
extension types are here for. And the browser test in
example/test/web_wasm_test.dart decodes through the real web binding and
asserts the RMS series arrives - the one leg neither of the others can see,
since a _copyOut that reads a series and then drops it on the floor compiles
and exports perfectly.
Changed #
- ABI 7 → 8. Additive:
wf_capture_scratch,wf_capture_scratch_frames,wf_capture_pcm_scratch,wf_capture_pcm_scratch_samplesandwf_capture_live. No existing signature changed.wf_capture_livereports sessions created and not yet destroyed, and exists for the same reasonwf_capture_feedis public: it is what makes the binding's ownership testable rather than asserted.
0.3.0 #
Capture keeps the audio, the pyramid carries loudness, and the native core actually loads on Android - which it never had.
Added #
- RMS through the pyramid. Peaks say how far the audio went; RMS says how
much of it there was.
WaveformPeaks.rms(level)andPeakWindow.rmsAt(i)expose it, and coarser levels combine children as the root of the mean of their squares so the value stays an RMS rather than an average of averages. Costs 50% more pyramid memory. - Capture keeps the audio. A second lock-free ring carries raw PCM
alongside the reduced frames, and
CaptureConfig.recordTostreams it to a 16-bit WAV. Separate rings because they run at 86/sec against 44,100/sec and a dropped frame is cosmetic where a dropped sample is a hole.CaptureSession.pcmDroppedis reported separately for that reason. CaptureSession.pauseandresume, which stop the device without touching the rings, the accumulator or the history, so a take continues rather than restarting.
Fixed: the native core never loaded on Android #
libmonowave.so was bundled correctly in every APK since M0 and failed to
dlopen on every one of them: cannot locate symbol "pow". miniaudio and
dr_mp3 both reference it, Apple's libSystem provides the math functions
implicitly, and Android and Linux do not - so libraries: ['m'] was missing
from the build hook.
M0 claimed Android was verified. What it actually verified was that the library was inside the APK, which is not the same as it loading, and the CI check inherited exactly that blind spot. It went unnoticed for five milestones because every test ran on the macOS host and the example's sample path is pure Dart. Running it on a real device is what found it.
Fixed: quiet recordings drew as a flat line #
The peaks painter scaled linearly, so a quiet room tone at about 1% of full
scale drew a two-pixel bar and read as broken. WaveformStyle.normalize scales
so the loudest moment fills the height, read from the coarsest mipmap level in
O(1) so it does not shift while panning. CompactBars already made the same
correction for the fixed-bar case.
Capture keeps the audio, not just the reduction #
Building a real voice-memo example surfaced a gap the milestones had missed:
stop() returned peaks, but the audio thread had only ever kept a reduction.
The PCM was gone, so a recording could be drawn and never trimmed or exported -
which is most of what a voice memo app does.
Added
- A second lock-free ring for raw PCM, alongside the one carrying reduced frames. The audio thread only ever copies into both; writing the file is the consumer's job, because file I/O on an audio callback is exactly the unbounded operation that produces a glitch.
CaptureConfig.recordTo, which streams the take to a 16-bit PCM WAV as it is captured. The header is written twice - once as a placeholder so the audio starts at a fixed offset, once at stop with the real sizes - so the file is streamed rather than assembled in memory.CaptureSession.pcmDropped, reported separately fromdropped. Losing a visualizer frame is cosmetic; losing audio is not.
The output is the same format the exporter reads, so a recording can be trimmed and exported without a second format in play.
The example is a voice memo app #
Replaced the developer gallery with a real product flow - idle, recording, review - with the waveform as the hero element throughout and the Android runtime microphone permission finally requested, closing the M3 gap.
0.2.0 #
Decode, capture, render and edit, on one C core across six targets. Web has everything but capture.
M5 - editing and export #
Added
- A non-destructive document: a list of regions referencing source ranges with gain and fades. Nothing decodes, copies or mutates audio; the source is untouched until an export reads it.
- Edits as values - trim, delete, split, gain, fade - in a sealed hierarchy, so an exporter switching over them fails to compile when a new kind is added.
EditHistorywith undo and redo over document snapshots rather than inverse operations, following monolens's reasoning: undo is cheap precisely because an edit is a value, and a fade has no inverse.previewPeaks, deriving the edited waveform from the source's peaks without decoding, so the display updates the instant an edit is applied.wf_export_wav: reads through the same decoders, seeks per region, applies gain and linear fades, writes 16-bit PCM WAV.- The Edit tab now trims, deletes, fades, undoes and exports.
Verified
An unedited document round-trips to byte-identical peaks, a trim exports exactly the selected range, a delete closes the gap, gain scales the output, and a fade reaches silence at its edges - all asserted by decoding the exported file back rather than by checking that a file appeared.
Notes
- Output is always WAV. An edit list is meant to reproduce the source exactly where it did not change it, and re-encoding to a lossy format would quietly break that.
- Fades are linear, not equal-power: these exist to take the click off an edit point, not to crossfade two takes, and linear is what makes the endpoints exactly 0 and 1.
previewPeaksdoes not show fades. A fade acts over samples and the finest level is 128 samples wide, so a typical fade is narrower than one bar.- Export is unavailable on web, which has no filesystem to write to.
M4 - selection, snapping and the zoom claim #
M1 already had the mipmap and the viewport maths. M4 is what you do once you are zoomed in, plus proof of the claim the pyramid exists to support.
Added
WaveformSelectionin sample space, so a range survives a zoom, a resize and a rotation without drifting. Immutable, which is also what will make undo cheap in M5.WaveformSnap-toZeroCrossingfinds a bucket whose extremes straddle zero,toQuietestfinds the least audible cut point in a window. Both work from peaks alone.- The Edit tab: pinch-zoom anywhere, drag to pan or to select depending on an explicit mode, and a selection overlay above the playhead layer.
Verified
A three-hour pyramid (476 million samples, 3.7 million pairs, 23 levels) resolves and reads every visible pair in 5.6 microseconds per frame through a full zoom sweep. A 60fps budget is 16,667 microseconds, so this is about 0.03% of it, and the cost is bounded by screen pixels rather than by file length. That is the entire payoff of the mipmap, and it is now a test with a threshold rather than a claim in prose.
Notes
- Snapping is bucket-accurate, not sample-accurate. At a 128-sample base that is about 3 ms - inaudible for a trim point, but worth stating plainly because "zero crossing" usually implies exactness. Sample-exact snapping would mean re-reading the source on every gesture.
- A drag cannot both navigate and select, so the Edit tab makes the mode explicit rather than guessing from a modifier. Pinch always zooms.
- Gesture state is captured at scale-start and updates are applied relative to it. Accumulating per-frame deltas drifts.
M3 - live capture #
The milestone where the constraints stop being preferences. Everything reachable from the audio callback is forbidden to allocate, take a lock, or call into a higher layer, and a missed deadline there is an audible glitch rather than a dropped frame.
Added
- The capture core, on vendored miniaudio. The audio thread accumulates a hop, reduces it to min/max/RMS, and publishes it through a lock-free single-producer/single-consumer ring. No PCM crosses into Dart: at a 512 sample hop that is about 516 bytes a second instead of 176 kB.
wf_capture_feedis public, and it is the entry point the device callback wraps. That is what makes the realtime path testable on every platform with no microphone, no permission prompt and exact timing - 17 tests drive it with synthetic PCM.- A preallocated history buffer, so
stop()returns peaks for the whole take even if the consumer never drained. Growing it from the audio thread is not an option, so it is sized up front and reports truncation past its cap. CaptureSession,CaptureScope,CaptureConfigin Dart. The scope is a ring over a preallocatedInt16Listand allocates nothing per frame.FakeCaptureSessioninpackage:monowave/testing.dart, emittable frame by frame so a host's visualizer can be tested without audio.- The Record tab, wired end to end: record, watch the bars, stop, and the captured peaks come back as the 64-byte summary a sender would upload.
Notes
- Dropped frames are surfaced rather than hidden. A full ring drops and counts; blocking the producer to wait for room would stall the audio device.
- Web capture is not implemented, and when it lands it will not use
miniaudio. M0's assumption that it would need
SharedArrayBufferturned out to be wrong twice over; see docs/20-concepts/90-architecture.md for what replaced it. - miniaudio must be compiled as Objective-C on Apple platforms, and
Language.objectiveConly adds-frameworkflags - clang picks the language from the file extension, so there is a one-linewf_miniaudio.mthat includes the.c. - The example declares the microphone permission on every platform, but Android also needs a runtime request that the example does not make yet. monowave deliberately does not request permissions: a headless package has no UI to explain why it is asking, and the host does.
M2 - the C decoder #
The first milestone where the single-C-core bet actually pays: real audio in, peaks out, over two completely different bindings.
Added
- The C decode core.
wf_decode_fileandwf_decode_memoryover vendored dr_wav, dr_mp3 and dr_flac, streaming a bucket at a time so peak memory is one bucket whether the input is a voice note or an audiobook. The pyramid is built in C and owned by C. MonowavePlatform.decodeFile/decodeBytes, implemented overdart:ffinatively and WASM on web. Native decodes run in a helper isolate and hand back only the pyramid's address, so nothing is copied or serialized across it.WaveformPeaks.fromLevelsanddispose(). On native, levels are typed-data views straight into native memory, kept alive by aNativeFinalizer, so a three-hour recording never touches the Dart heap.- Fixtures and the determinism check. Six synthesized WAVs, each targeting a
specific way a reduction can be wrong, hashed into digests that
dart testasserts on every OS andtool/verify_wasm.mjsasserts against the WASM build.
Verified
All six fixtures decode to byte-identical pyramids in the WASM build and the native build. That is the property the architecture exists to guarantee, and it now has a test rather than an argument.
Notes
- AAC/M4A is not supported, and cannot be without a platform decoder or a much heavier dependency. It is tolerable because the voice-note path never decodes: the sender computes peaks at record time and ships them as metadata.
- Determinism is asserted on WAV and FLAC, whose paths are integer-exact. MP3 decodes through floating point, where the last bit can legitimately differ between targets.
wf_peaks_lengthreturns adouble. Anint64would surface to JavaScript as aBigInt, and 2^53 samples is over a thousand years of audio.- The WASM module imports three WASI file-descriptor functions that libc links in even with the decoders' file halves compiled out. Both bindings stub them with ENOSYS; they are unreachable from the decode path.
M1 - model, codecs and the gallery #
Everything here is pure Dart. There is still no decoder, and none of it needs one: a sender computes peaks at record time, so the common path never decodes anything.
Added
WaveformPeaks- a min/max mipmap pyramid. Level 0 is the finest resolution held; each level above is built by min-of-mins and max-of-maxes, so a coarse level provably bounds the level below it and zooming is exact rather than approximate. Reduction is never an average.WaveformViewportandPeakWindow- which part of the audio is on screen and at what zoom, resolved to a level and handed to a painter as a zero-copy window. Zooming anchors on a focus point so a pinch feels attached to the audio rather than to the widget.WaveformTimeline- the whole of monowave's relationship with a player:Durationto sample and back, with no player dependency.WaveformDat- the BBCaudiowaveformbinary format, versions 1 and 2 at 8 or 16 bits, so peaks can be precomputed server-side by the standard tool.CompactBars- a 64-byte voice-note summary with normalization and a dBFS curve, plusfromAmplitudesfor the live-capture path.- The example is now a monokit v2.0.0 gallery. It composes rather than
duplicates: monokit's own
MonoWaveformandMonoVoiceNoterender the fixed-bar case fromCompactBars, and a reference painter covers what they cannot - true min/max asymmetry and a viewport that zooms.
M0 - toolchain spike #
Nothing was published from this milestone; it existed to settle how the native side is built before any of it was written.
Added
- Repo scaffold following the monokit/monolens house layout: lefthook and
commitlint over conventional commits,
flutter_lints, MIT license, generated (never committed) fixtures. hook/build.dartcompilingsrc/into a code asset viapackage:hooksandpackage:native_toolchain_c, replacing the per-platform CMake, podspec and Gradle scaffolding a classic FFI plugin would carry.src/wf_probe.c- the M0 probe.wf_reduce_minmaxis the real peak kernel in miniature, so the spike proves pointer passing and not just linking.MonowavePlatform, the mockable seam, with its two real implementations selected by conditional import:dart:ffion the five native targets, anddart:js_interopover WASM on web.tool/build_wasm.sh, producing the committedassets/monowave.wasm. It carries workarounds for two bugs in Homebrew's emscripten 6.0.4 bottle so a fresh clone builds without manual setup.FakeMonowavePlatforminpackage:monowave/testing.dart.
Verified
The same C source, reached over FFI natively and over WASM on web, returns identical results on every target checked: macOS (host and app), iOS simulator at runtime, Android for all three ABIs, and web at runtime in a browser. Linux and Windows are covered by the CI matrix rather than locally.
Notes
- Tests are
package:testrather thanflutter_test. A headless package has no widget tree to bind, andflutter_testpinsmeta 1.18.0from the SDK, which collides with the hook packages. hooksandnative_toolchain_care held one patch below latest: both moved tometa ^1.19.0, which Flutter stable's pinnedmeta 1.18.0cannot satisfy.