Abstract

Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.

SpanSynth-Edit model overview. Panel (a) shows contextual audio, target audio latents, and frame-aligned MIDI conditioning for the flow-matching model. Panel (b) shows instrument-labeled notes encoded as unordered MIDI Span events and pooled into frame features.

Audio demos

Listen alongside the MIDI and reference audio.

EARLY DEMO

Editing