BARS: Beat-Adaptive Rap Synthesis via Flow Matching, accepted to ISMIR 2026

Excited to share that our work at BandLab Technologies, “BARS: Beat-Adaptive Rap Synthesis via Flow Matching,” has been accepted to ISMIR 2026!

Given lyrics, a beat track, and a short reference voice, BARS generates rap vocals that perform the lyrics over the beat using the timbre of the reference voice.

Unlike prior autoregressive approaches, BARS uses a non-autoregressive Flow Matching architecture, achieving a real-time factor of 0.027 on an NVIDIA RTX 4090 GPU. That’s about 1.6 seconds to generate one minute of rap, end-to-end. It also requires no explicit duration prediction or phoneme-level alignment during training, and allows independent control over rap tempo for flexible-length generation while maintaining rhythmic coherence with the beat.

Huge thanks to my wonderful co-authors Héctor Martel - 何可拓 and En Yan Koh for their help along the way.

The full paper will be available later. For now, we’ve put together a demo page where you can listen to examples generated by BARS.

See you at ISMIR 2026!




Enjoy Reading This Article?

Here are some more articles you might like to read next: