LLM-Guided Affective Music Generation via Prompt-to-Parameter Mapping and Frequency-Aware Evaluation
DOI:
https://doi.org/10.6911/Keywords:
Text-To-Music Generation; Affective Music Generation; Large Language Models; Prompt-To-Parameter Mapping; Frequency-Aware Evaluation.Abstract
Recent text-to-music generation models have enabled users to compose music from natural language description; they typically regard emotional prompts as a black-box conditioning signal, which is hard to interpret, control, and assess. In this paper, we present an LLM-guided affective music creation framework featuring prompt-to-parameter transformation and frequency-aware evaluation. Our approach begins by taking a handwritten emotional cue and applying a large language model as affective mapper that converts subjective language into structured music parameters such as tempo, dynamics, harmony, texture density, register, instrumentation and spectral targets. These parameters are then input into a downstream model that generates music. For music generation evaluation, we propose a multi-dimensional protocol on the aspects of audio fidelity, text-music alignment, and frequency-domain consistency. We evaluate 40 affective prompts, and our method achieves 0.536 of CLAP alignment, 4.31 of FAD, and 0.792 of spectral consistency, while direct prompt generation achieves 0.412, 5.87, and 0.641, respectively. Subjective listening results also reveal that our method results in higher overall quality, emotional consistency and listening comfort. These results show that the explicit LLM-based affective mapping offers a more interpretable and controllable approach to emotional music generation.
Downloads
References
[1] Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., Sharifi, M., Frank, C., & Zeghidour, N. (2023). MusicLM: Generating music from text. arXiv preprint arXiv:2301.11325. https://doi.org/10.48550/arXiv.2301.11325.
[2] Aljanaki, A., Yang, Y.-H., & Soleymani, M. (2017). Developing a benchmark for emotional analysis of music. PLOS ONE, 12(3), e0173392. https://doi.org/10.1371/journal.pone.0173392
[3] Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., & Défossez, A. (2023). Simple and controllable music generation. Advances in Neural Information Processing Systems, 36.
[4] Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., & Sutskever, I. (2020). Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341. https://doi.org/10.48550/arXiv.2005.00341
[5] Elizalde, B., Deshmukh, S., Al Ismail, M., & Wang, H. (2023). CLAP: Learning audio concepts from natural language supervision. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, 1–5. https://doi.org/10.1109/ICASSP40776.2023.10096654
[6] Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., & Ellis, D. P. W. (2022). MuLan: A joint embedding of music audio and natural language. Proceedings of the International Society for Music Information Retrieval Conference.
[7] Kilgour, K., Zuluaga, M., Roblek, D., & Sharifi, M. (2019). Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms. Proceedings of Interspeech, 2350–2354. https://doi.org/10.21437/Interspeech.2019-2680
[8] Liu, H., Chen, Z., Yuan, Y., Mei, X., Liu, X., Mandic, D., Wang, W., & Plumbley, M. D. (2023). AudioLDM: Text-to-audio generation with latent diffusion models. Proceedings of the International Conference on Machine Learning.
[9] McFee, B., Raffel, C., Liang, D., Ellis, D. P. W., McVicar, M., Battenberg, E., & Nieto, O. (2015). librosa: Audio and music signal analysis in python. Proceedings of the Python in Science Conference, 18–25. https://doi.org/10.25080/Majora-7b98e3ed-003.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 World Scientific Research Journal

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




