JNMR 2026

Change is Key: A Generative Framework for Controllable Musical Modulations

Anonymous Authors

Abstract

The ability to transition between two tonalities, or to modulate, is a vital compositional tool for shaping emotional contour and musical tension. However, controlling the generation of key changes in a meaningful and musically coherent manner is a significant challenge for AI-driven music systems, owing to the lack of large scale annotated modulation datasets. To overcome this data scarcity problem, we formulate musical modulation as a generative infilling task that enables models to learn from vast amounts of unannotated data via self-supervised training. We evaluate this framework through a purpose-built LLaMA-based infilling model which supports both an autonomous generation mode and a rule-guided model conditioned on harmonic progressions. To assess the modulatory mechanics without relying on ground-truth annotations, we design a multi-dimensional, reference-free evaluation suite quantifying stylistic coherence, rhythmic continuity, voice-leading smoothness, and harmonic tension. We validate these metrics through a subjective listening test to demonstrate a meaningful correlation with human aesthetic perception. Furthermore, zero-shot evaluations on both human-composed reconstructions and forced modulations across varying semitone shifts show that our autonomous model mitigates harmonic shock and maintains stylistic cohesion, even when navigating distant tonal intervals.

Highlights

The following two examples illustrate the same modulation, from A♭ minor to A minor, as composed by a human and as generated autonomously by our model. They serve as a representative showcase of what our framework can achieve.

Human-Composed (Ground Truth) – A♭ minor to A minor
Commentary: The piece modulates from A♭ minor to A minor. The original human-composed bridge from the POP909 dataset.
Autonomous Model – A♭ minor to A minor
Commentary: The piece modulates from A♭ minor to A minor. This example is particularly noteworthy because the model not only generates harmonically coherent and well-resolved voice leading, but also introduces rhythmic novelty within the transition window, making it one of the most musically interesting outputs produced by our framework.

Modulation Examples

Each example shows five versions of the same modulation window: the human-composed reference (Ground Truth) and four generated variants. Use the tabs to navigate between examples.

Note: In all clips, the generated modulation window spans from 00:20 to 00:30. The context preceding and succeeding it are the same as the original human composition across all methods. The Pianoteq chords may not always be accurate. Please use headphones for the best listening experience.