Lunar Boom Learning

Sources

Course reference

Sources and citations

Each source remains connected to the lesson that uses it, even where a citation appears with different contextual notes.

Chapter 5: Evaluating, Releasing, and Governing the Model (121)
  1. 45 CFR 46: Protection of Human Subjects

    Office for Human Research Protections · United States Department of Health and Human Services · regulation

    Official United States regulations for covered human-subject research. Applicability depends on funding, institution, jurisdiction, and study characteristics.

    Cited in Human Listening Tests
  2. ACE-Step: A Step Towards Music Generation Foundation Model

    Junmin Gong, Sean Zhao, Sen Wang, Shengyuan Xu, and Jing Guo · arXiv · 2025 · paper

    Primary research on a compressed diffusion and transformer music foundation model with multi-minute generation, temporal repainting, lyric editing, remixing, and specialized adaptation tasks.

    Cited in Where AI Music Models Are Heading
  3. Adapting Frechet Audio Distance for Generative Music Evaluation

    Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Emmanouilidou · IEEE ICASSP · 2024 · paper

    Music-specific analysis of FAD sensitivity to sample count, embedding choice, reference-set quality, and perceptual correlation.

    Cited in What Makes Generated Music Good?
  4. Adapting Fréchet Audio Distance for Generative Music Evaluation

    Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Emmanouilidou · IEEE ICASSP · 2024 · paper

    Music-specific study of FAD sample-size bias, audio-embedding choice, reference-set quality, and perceptual correlation.

    Cited in Automated Evaluation
  5. AI Risk Management Framework Playbook: Manage

    NIST · National Institute of Standards and Technology · government-guidance

    Official guidance for deciding whether an AI system achieves its intended purpose and whether development or deployment should proceed.

    Cited in What Makes Generated Music Good?
  6. AI Risk Management Framework Playbook: Manage

    NIST · National Institute of Standards and Technology · government-guidance

    Official guidance on post-deployment monitoring, external feedback, performance degradation, unexpected behavior, incidents, and risk response.

    Cited in Turning a Model Into a Product
  7. AI Risk Management Framework Playbook: Map

    NIST · National Institute of Standards and Technology · government-guidance

    Official guidance for documenting intended purpose, users, deployment context, expectations, and impacts before selecting evaluation criteria.

    Cited in What Makes Generated Music Good?
  8. AI Risk Management Framework Playbook: Measure

    NIST · National Institute of Standards and Technology · government-guidance

    Official guidance stating that evaluation approaches and metrics depend on purpose, audience, needs, and context, and should test whether a system is fit for purpose.

    Cited in What Makes Generated Music Good?
  9. AI Risk Management Framework Playbook: Measure

    NIST · National Institute of Standards and Technology · government-guidance

    Official guidance emphasising that evaluation methods and metrics should be selected according to purpose, audience, context, and intended use.

    Cited in Automated Evaluation
  10. AI Risk Management Framework Playbook: Measure

    NIST · National Institute of Standards and Technology · government-guidance

    Official guidance on selecting evaluation methods according to purpose, audience, context, intended use, and decision requirements.

    Cited in Human Listening Tests
  11. Algorithms to Measure Audio Programme Loudness and True-Peak Audio Level

    ITU Radiocommunication Sector · International Telecommunication Union · 2023 · standard

    Official recommendation specifying algorithms for programme loudness and true-peak measurement.

    Cited in Human Listening Tests
  12. Amazon SQS Visibility Timeout

    AWS · Amazon Web Services · documentation

    Official explanation of message invisibility during processing and redelivery when processing is not acknowledged before the timeout.

    Cited in Turning a Model Into a Product
  13. An Early Look at the Possibilities as We Experiment with AI and Music

    Lyor Cohen and Toni Reid · YouTube · 2023 · official-product-announcement

    Official announcement of Dream Track, developed with participating artists who chose to collaborate and initially offered to selected creators.

    Cited in Where AI Music Models Are Heading
  14. An Industrial-Strength Audio Search Algorithm

    Avery Li-Chun Wang · International Society for Music Information Retrieval · 2003 · paper

    Foundational description of a scalable spectral-landmark audio-fingerprinting system for identifying recordings from short, noisy, or compressed excerpts.

    Cited in Similarity, Memorisation, and Attribution
  15. Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls

    Liwei Lin, Gus Xia, Yixiao Zhang, and Junyan Jiang · arXiv · 2024 · paper

    Primary research adapting MusicGen for inpainting, score-conditioned arrangement, and track-conditioned refinement.

    Cited in Where AI Music Models Are Heading
  16. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile

    NIST · National Institute of Standards and Technology · 2024 · government-report

    Official cross-sector guidance for managing generative-AI risks, including governance, pre-deployment testing, content provenance, monitoring, and incident disclosure.

    Cited in Turning a Model Into a Product
  17. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile

    NIST · National Institute of Standards and Technology · 2024 · government-guidance

    Official cross-sector guidance for identifying, measuring, managing, and governing generative-AI risks across development and deployment.

    Cited in Where AI Music Models Are Heading
  18. Asynchronous Inference

    AWS · Amazon Web Services · documentation

    Official example of queued asynchronous model serving for long processing times and large payloads, including scale-to-zero support.

    Cited in Turning a Model Into a Product
  19. audiocraft.models.musicgen API Documentation

    Meta AI · Meta AudioCraft · documentation

    Official MusicGen generation API combining the required model components and exposing text and melody generation methods.

    Cited in Turning a Model Into a Product
  20. AudioGen: Textually Guided Audio Generation

    Felix Kreuk et al. · International Conference on Learning Representations · 2023 · paper

    Original AudioGen paper using objective and subjective evaluation for text-conditioned audio generation.

    Cited in Automated Evaluation
  21. AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

    Haohe Liu et al. · International Conference on Machine Learning · 2023 · paper

    Original AudioLDM paper using CLAP representations and objective and subjective text-to-audio evaluation.

    Cited in Automated Evaluation
  22. Automatic Evaluation of Speaker Similarity

    Deja Kamil, Sanchez Ariadna, Roth Julian, and Cotescu Marius · arXiv · 2022 · preprint

    Studies automatic speaker-similarity prediction using speaker embeddings and comparison with human perceptual ratings.

    Cited in Similarity, Memorisation, and Attribution
  23. Autoscale an Asynchronous Endpoint

    AWS · Amazon Web Services · documentation

    Official guidance for autoscaling queued inference endpoints, including scaling to zero.

    Cited in Turning a Model Into a Product
  24. Benchmarking Music Generation Models and Metrics via Human Preference Studies

    Florian Grötschla, Ahmet Solak, Luca A. Lanzendörfer, and Roger Wattenhofer · arXiv · 2025 · paper

    Large human-preference benchmark using generated songs from multiple systems, illustrating the difficulty of aligning automated music metrics with human judgement.

    Cited in Where AI Music Models Are Heading
  25. C2PA and Content Credentials Explainer, Version 2.4

    C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-guidance

    Official explanation of Content Credentials, provenance, asset bindings, ingredients, AI disclosure, durable credentials, trust, privacy, and the limits of provenance assertions.

    Cited in Similarity, Memorisation, and Attribution
  26. C2PA Technical Specification, Version 2.4

    C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-standard

    Current technical specification for cryptographically bound Content Credentials, including audio actions, ingredients, AI models, datasets, prompts, seeds, and generative-system asset types.

    Cited in Similarity, Memorisation, and Attribution
  27. C2PA Technical Specification, Version 2.4

    C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-standard

    Current technical specification for cryptographically bound Content Credentials, asset actions, ingredients, audit logs, and AI and ML asset types.

    Cited in Turning a Model Into a Product
  28. C2PA Technical Specification, Version 2.4

    C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-standard

    Current technical specification for Content Credentials, asset actions, ingredients, AI and ML asset types, and provenance assertions.

    Cited in Where AI Music Models Are Heading
  29. Capability: Unit Economics

    FinOps Foundation · FinOps Foundation · industry-guidance

    Official FinOps guidance connecting cloud costs to product units, engineering decisions, pricing, and business outcomes.

    Cited in Turning a Model Into a Product
  30. CLAP: Learning Audio Concepts From Natural Language Supervision

    Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · IEEE ICASSP · 2023 · paper

    Original CLAP paper describing contrastive audio and language encoders trained to produce a joint embedding space.

    Cited in What Makes Generated Music Good?
  31. CLAP: Learning Audio Concepts From Natural Language Supervision

    Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · IEEE ICASSP · 2023 · paper

    Original CLAP paper describing contrastive audio and text encoders trained to produce a shared multimodal representation.

    Cited in Automated Evaluation
  32. Cloud Cost Allocation

    FinOps Foundation · FinOps Foundation · industry-guidance

    Official FinOps guidance for assigning cloud costs to users, products, teams, projects, and other accountable groupings.

    Cited in Turning a Model Into a Product
  33. Coarse Parallel Processing Using a Work Queue

    Kubernetes contributors · Kubernetes · documentation

    Official example in which parallel workers claim units of work from a task queue.

    Cited in Turning a Model Into a Product
  34. CROWDMOS: An Approach for Crowdsourcing Mean Opinion Score Studies

    Flavio Ribeiro, Dinei Florêncio, Cha Zhang, and Mike Seltzer · IEEE ICASSP · 2011 · paper

    Original crowdsourced subjective-audio study proposing methods for identifying inaccurate responses and obtaining repeatable quality judgements outside a laboratory.

    Cited in Human Listening Tests
  35. Datasheets for Datasets

    Timnit Gebru et al. · Communications of the ACM · 2021 · paper

    Original proposal for dataset documentation covering motivation, composition, collection, preprocessing, uses, distribution, maintenance, and limitations.

    Cited in Similarity, Memorisation, and Attribution
  36. Datasheets for Datasets

    Timnit Gebru et al. · Communications of the ACM · 2021 · paper

    Original proposal for documenting dataset motivation, composition, collection, processing, uses, maintenance, and limitations.

    Cited in Where AI Music Models Are Heading
  37. Download and Upload Objects with Presigned URLs

    AWS · Amazon Web Services · documentation

    Official guidance for time-limited upload and download access without exposing general storage credentials.

    Cited in Turning a Model Into a Product
  38. ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

    Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck · Interspeech · 2020 · paper

    Primary paper introducing a widely used speaker-embedding architecture for automatic speaker verification.

    Cited in Similarity, Memorisation, and Attribution
  39. EnCodec Documentation

    Meta AI · Meta AudioCraft · documentation

    Official AudioCraft documentation for training and evaluating EnCodec.

    Cited in Automated Evaluation
  40. Extracting Training Data from Large Language Models

    Nicholas Carlini et al. · USENIX Security Symposium · 2021 · paper

    Foundational demonstration that model querying can recover identifiable training examples. Included for the distinction between memorisation and extractability.

    Cited in Similarity, Memorisation, and Attribution
  41. Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

    Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · Interspeech · 2019 · paper

    Original Fréchet Audio Distance paper comparing distribution-level audio embeddings and human perception of audio distortions.

    Cited in What Makes Generated Music Good?
  42. Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

    Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · Interspeech · 2019 · paper

    Original Fréchet Audio Distance paper introducing distributional comparison of audio embeddings and validating it against artificial distortions and human perception.

    Cited in Automated Evaluation
  43. General Methods for the Subjective Assessment of Sound Quality

    ITU Radiocommunication Sector · International Telecommunication Union · 2019 · standard

    Official recommendation covering general sound-quality assessment, grading scales, comparison tests, random presentation order, listening conditions, statistical analysis, and reporting.

    Cited in Human Listening Tests
  44. Guidance for Artificial Intelligence and Machine Learning

    C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-guidance

    Official guidance for applying Content Credentials to datasets, models, software, and media used during training and inference.

    Cited in Turning a Model Into a Product
  45. Guidance for Artificial Intelligence and Machine Learning, Version 2.4

    C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-guidance

    Official guidance for applying Content Credentials to AI datasets, software, models, fine-tuning, inference, outputs, versioning, and release.

    Cited in Where AI Music Models Are Heading
  46. Guidance for the Selection of the Most Appropriate ITU-R Recommendation for Subjective Assessment of Sound Quality

    ITU Radiocommunication Sector · International Telecommunication Union · standard

    Official guidance for selecting a subjective audio-assessment method according to the test purpose and expected impairment.

    Cited in Human Listening Tests
  47. High Fidelity Neural Audio Compression

    Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · Transactions on Machine Learning Research · 2023 · paper

    Original EnCodec paper combining reconstruction objectives, perceptual losses, ablations, and MUSHRA listening tests for music and other audio domains.

    Cited in Automated Evaluation
  48. How Robust is Audio Watermarking in Generative AI Models?

    Yuxin Wen et al. · arXiv · 2025 · preprint

    Evaluates recent audio-watermarking systems against a broad set of removal attacks and highlights the need to test claimed robustness under realistic transformations.

    Cited in Similarity, Memorisation, and Attribution
  49. inference_mode

    PyTorch contributors · PyTorch · documentation

    Official documentation for disabling gradient-related tracking and selected autograd overhead during inference.

    Cited in Turning a Model Into a Product
  50. Informed Consent FAQs

    Office for Human Research Protections · United States Department of Health and Human Services · government-guidance

    Official guidance on legally effective and documented informed consent for research covered by HHS requirements.

    Cited in Human Listening Tests
  51. InvokeEndpointAsync

    AWS · Amazon Web Services · api-documentation

    Official API showing that asynchronous requests are enqueued and return result-location information rather than the inference result.

    Cited in Turning a Model Into a Product
  52. ITU-R BS.1116: Methods for the Subjective Assessment of Small Impairments in Audio Systems

    ITU Radiocommunication Sector · International Telecommunication Union · 2015 · standard

    Official recommendation covering rigorous experimental design, listening conditions, subject selection, and reporting for critical audio assessment.

    Cited in What Makes Generated Music Good?
  53. ITU-R BS.1283: Subjective Assessment of Sound Quality

    ITU Radiocommunication Sector · International Telecommunication Union · 1997 · standard

    Official guide explaining that subjective audio-assessment methods should be selected according to the purpose and scale of the expected impairments.

    Cited in What Makes Generated Music Good?
  54. ITU-R BS.1534: Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems

    ITU Radiocommunication Sector · International Telecommunication Union · 2015 · standard

    Official recommendation defining the MUSHRA method and controlled subjective assessment principles for intermediate audio quality.

    Cited in What Makes Generated Music Good?
  55. Jobs

    Kubernetes contributors · Kubernetes · documentation

    Official documentation for run-to-completion workloads, retries, parallelism, completion tracking, and failure policies.

    Cited in Turning a Model Into a Product
  56. Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

    Or Tal, Alon Ziv, Itai Gat, Felix Kreuk, and Yossi Adi · IEEE ICASSP · 2024 · paper

    Primary JASCO paper combining text with temporally aligned chord, melody, drum, and full-mix controls.

    Cited in Where AI Music Models Are Heading
  57. librosa Feature Extraction Documentation

    librosa contributors · librosa · documentation

    Official documentation for chroma, spectral, mel, tonal, onset, rhythm, and related audio features.

    Cited in Automated Evaluation
  58. librosa.beat.beat_track

    librosa contributors · librosa · documentation

    Official documentation for dynamic-programming beat tracking using onset strength and tempo estimation.

    Cited in Automated Evaluation
  59. Live Music Models

    Lyria Team et al. · arXiv · 2025 · paper

    Primary research introducing continuous real-time music streams with synchronized control and describing Magenta RealTime and Lyria RealTime.

    Cited in Where AI Music Models Are Heading
  60. Long-form Music Generation with Latent Diffusion

    Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · arXiv · 2024 · paper

    Primary research reporting latent-diffusion music generation over long temporal contexts, with generated tracks of up to four minutes and forty-five seconds.

    Cited in Where AI Music Models Are Heading
  61. Lyria 3

    Google DeepMind · Google DeepMind · official-model-documentation

    Official documentation describing Lyria 3, multimodal image conditioning, technical controls, and generation of cohesive tracks up to three minutes.

    Cited in Where AI Music Models Are Heading
  62. Lyria RealTime

    Google DeepMind · Google DeepMind · official-model-documentation

    Official documentation describing real-time interactive music creation, control, and performance.

    Cited in Where AI Music Models Are Heading
  63. Managing the Lifecycle of Objects

    AWS · Amazon Web Services · documentation

    Official documentation for transitioning objects to lower-cost classes and expiring objects through lifecycle rules.

    Cited in Turning a Model Into a Product
  64. MelodySim: Measuring Melody-Aware Music Similarity for Plagiarism Detection

    Tianze Lu et al. · arXiv · 2025 · preprint

    Emerging research on melody-aware music similarity under transformations including note splitting, arpeggiation, track dropout, and re-instrumentation. Included as a developing method rather than an established universal standard.

    Cited in Similarity, Memorisation, and Attribution
  65. MERT: Acoustic Music Understanding Model with Large-Scale Self-Supervised Training

    Yizhi Li et al. · International Conference on Learning Representations · 2024 · paper

    Original MERT paper describing a music-specialised self-supervised representation evaluated across fourteen music-understanding tasks.

    Cited in Automated Evaluation
  66. Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems

    ITU Radiocommunication Sector · International Telecommunication Union · 2023 · standard

    Official recommendation defining the MUSHRA method with multiple stimuli, hidden reference, anchors, controlled presentation, and subjective quality analysis.

    Cited in Human Listening Tests
  67. Methods for the Subjective Assessment of Small Impairments in Audio Systems

    ITU Radiocommunication Sector · International Telecommunication Union · 2015 · standard

    Official recommendation for rigorous subjective evaluation of small audio impairments, including listening conditions, subject selection, training, test design, analysis, and reporting.

    Cited in Human Listening Tests
  68. mir_eval Documentation

    Colin Raffel et al. · mir_eval contributors · documentation

    Official documentation for standardised evaluation functions used in music-information-retrieval research.

    Cited in Automated Evaluation
  69. mir_eval: A Transparent Implementation of Common MIR Metrics

    Colin Raffel et al. · International Society for Music Information Retrieval · 2014 · paper

    Primary paper describing transparent implementations for beat, chord, melody, onset, pattern, segmentation, and other music-information-retrieval metrics.

    Cited in Automated Evaluation
  70. mir_eval: A Transparent Implementation of Common MIR Metrics

    Colin Raffel et al. · International Society for Music Information Retrieval · 2014 · paper

    Primary source for transparent evaluation implementations covering melody, transcription, segmentation, rhythm, and other music-information-retrieval tasks.

    Cited in Similarity, Memorisation, and Attribution
  71. mir_eval.melody Documentation

    mir_eval contributors · mir_eval · documentation

    Official documentation for evaluating voicing and pitch contours in melody extraction.

    Cited in Similarity, Memorisation, and Attribution
  72. ML Model Registry

    MLflow contributors · MLflow · documentation

    Official documentation for registered model versions, aliases, tags, source runs, and deployment-oriented version management.

    Cited in Turning a Model Into a Product
  73. MLflow Model Serving

    MLflow contributors · MLflow · documentation

    Official guidance on model signatures, environment capture, registry aliases, request validation, metadata, and serving.

    Cited in Turning a Model Into a Product
  74. Model Cards for Model Reporting

    Margaret Mitchell et al. · ACM Conference on Fairness, Accountability, and Transparency · 2019 · paper

    Original proposal for model cards documenting intended uses, evaluation procedures, performance characteristics, limitations, and ethical considerations.

    Cited in Similarity, Memorisation, and Attribution
  75. Model Cards for Model Reporting

    Margaret Mitchell et al. · ACM Conference on Fairness, Accountability, and Transparency · 2019 · paper

    Original proposal for structured model documentation covering intended use, evaluation, limitations, and relevant context.

    Cited in Where AI Music Models Are Heading
  76. Model Hosting FAQs

    AWS · Amazon Web Services · documentation

    Official comparison of real-time, asynchronous, serverless, and batch inference patterns.

    Cited in Turning a Model Into a Product
  77. Model Management

    NVIDIA · NVIDIA · documentation

    Official guidance on loading, unloading, adding, and removing model versions while serving traffic.

    Cited in Turning a Model Into a Product
  78. Model Repository

    NVIDIA · NVIDIA · documentation

    Official model-repository layout and version-policy documentation.

    Cited in Turning a Model Into a Product
  79. MuLan: A Joint Embedding of Music Audio and Natural Language

    Qingqing Huang et al. · International Society for Music Information Retrieval · 2022 · paper

    Original MuLan paper linking music audio with unconstrained natural-language descriptions for retrieval and zero-shot music understanding.

    Cited in What Makes Generated Music Good?
  80. MuLan: A Joint Embedding of Music Audio and Natural Language

    Qingqing Huang et al. · International Society for Music Information Retrieval · 2022 · paper

    Original MuLan paper linking music audio with unconstrained natural-language descriptions for retrieval and zero-shot applications.

    Cited in Automated Evaluation
  81. Music AI Sandbox, Now with New Features and Broader Access

    Google DeepMind · Google DeepMind · 2025 · official-product-announcement

    Official description of Lyria 2, Lyria RealTime, musician collaboration, continuous control, and SynthID watermarking.

    Cited in Where AI Music Models Are Heading
  82. MusicGen Documentation

    Meta AI · Meta AudioCraft · documentation

    Official overview of MusicGen models, EnCodec tokenization, pretrained checkpoints, and inference use.

    Cited in Turning a Model Into a Product
  83. MusicGen-Stem: Multi-stem Music Generation and Edition Through Autoregressive Modeling

    Simon Rouard, Robin San Roman, Yossi Adi, and Axel Roebel · IEEE ICASSP · 2025 · paper

    Primary research on parallel bass, drum, and other-stem generation, editing, and iterative source-conditioned composition.

    Cited in Where AI Music Models Are Heading
  84. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original MusicLM paper reporting separate evaluation of audio quality and text adherence and introducing MusicCaps.

    Cited in What Makes Generated Music Good?
  85. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · dataset-paper

    Primary source for MusicCaps, described as approximately 5,500 music-text pairs with rich expert-written descriptions.

    Cited in What Makes Generated Music Good?
  86. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original MusicLM paper using automated and human measures of audio quality and text adherence and introducing the MusicCaps benchmark.

    Cited in Automated Evaluation
  87. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · dataset-paper

    Primary source for MusicCaps, a benchmark of approximately 5,500 music-text pairs with detailed human-written descriptions.

    Cited in Automated Evaluation
  88. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original MusicLM paper describing a pairwise human test with two ten-second clips, a caption, graded A/B preference responses, and instructions focused specifically on text adherence.

    Cited in Human Listening Tests
  89. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original MusicLM paper containing a training-data memorisation analysis based on known training examples, prefixes, and continuation similarity.

    Cited in Similarity, Memorisation, and Attribution
  90. MusicRL: Aligning Music Generation to Human Preferences

    Geoffrey Cideron et al. · Proceedings of Machine Learning Research · 2024 · paper

    Peer-reviewed ICML paper describing 300,000 pairwise music preferences and showing that text adherence and audio quality explain only part of human musical preference.

    Cited in Human Listening Tests
  91. MusicRL: Aligning Music Generation to Human Preferences

    Geoffrey Cideron et al. · Proceedings of Machine Learning Research · 2024 · paper

    Peer-reviewed research using reward models and 300,000 pairwise preferences to align a MusicLM-based generator.

    Cited in Where AI Music Models Are Heading
  92. Observability Primer

    OpenTelemetry contributors · OpenTelemetry · documentation

    Official introduction to observability and instrumentation for understanding distributed systems.

    Cited in Turning a Model Into a Product
  93. OWASP Top 10 for Large Language Model Applications

    OWASP contributors · OWASP · security-guidance

    Community security guidance covering prompt injection, unsafe downstream output handling, excessive permissions, sensitive information, and related application risks.

    Cited in Turning a Model Into a Product
  94. Proactive Detection of Voice Cloning with Localized Watermarking

    Robin San Roman et al. · International Conference on Machine Learning · 2024 · paper

    Introduces AudioSeal, a neural audio watermark with localised detection designed for identifying AI-generated speech regions.

    Cited in Similarity, Memorisation, and Attribution
  95. Prompt Registry

    MLflow contributors · MLflow · documentation

    Official documentation for immutable prompt versions, aliases, comparison, reuse, A/B testing, and rollback.

    Cited in Turning a Model Into a Product
  96. Quantifying Memorization Across Neural Language Models

    Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang · International Conference on Learning Representations · 2023 · paper

    Primary research connecting discoverable memorisation with model capacity, duplicated examples, and prompting context. MusicLM and MusicGen adapt related methodology to music.

    Cited in Similarity, Memorisation, and Attribution
  97. Ragged Batching

    NVIDIA · NVIDIA · documentation

    Official guidance on batching requests with variable input shapes and the padding trade-off.

    Cited in Turning a Model Into a Product
  98. Rank Analysis of Incomplete Block Designs: The Method of Paired Comparisons

    Ralph Allan Bradley and Milton E. Terry · Biometrika · 1952 · paper

    Foundational paper introducing the Bradley-Terry method for estimating rankings from pairwise comparison outcomes.

    Cited in Human Listening Tests
  99. Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

    Baisen Wang, Chenxi Bao, and Qisong Han · arXiv · 2026 · preprint

    Recent preprint exploring streaming latent generation, consistency distillation, and dynamic control for low-latency interaction. Included as emerging research rather than settled evidence.

    Cited in Where AI Music Models Are Heading
  100. Secure Software Development Practices for Generative AI and Dual-Use Foundation Models

    NIST · National Institute of Standards and Technology · 2024 · government-report

    Official secure-development profile for producers and acquirers of AI models and AI systems.

    Cited in Turning a Model Into a Product
  101. Signals

    OpenTelemetry contributors · OpenTelemetry · documentation

    Official definitions of traces, metrics, logs, baggage, and profiles.

    Cited in Turning a Model Into a Product
  102. Simple and Controllable Music Generation

    Jade Copet et al. · NeurIPS · 2023 · paper

    Original MusicGen paper combining automated evaluation, human ratings of audio quality and text relevance, ablations, and memorisation testing.

    Cited in What Makes Generated Music Good?
  103. Simple and Controllable Music Generation

    Jade Copet et al. · NeurIPS · 2023 · paper

    Original MusicGen paper combining text-alignment, audio-quality, human, diversity, ablation, and memorisation evaluations.

    Cited in Automated Evaluation
  104. Simple and Controllable Music Generation

    Jade Copet et al. · NeurIPS · 2023 · paper

    Original MusicGen paper reporting separate human ratings for overall perceptual quality and text relevance, crowdsourced response filtering, several ratings per sample, confidence intervals, and loudness normalisation.

    Cited in Human Listening Tests
  105. Simple and Controllable Music Generation

    Jade Copet et al. · NeurIPS · 2023 · paper

    Original MusicGen paper describing token-based memorisation analysis, training-example continuations, similarity thresholds, and human and automated evaluation.

    Cited in Similarity, Memorisation, and Attribution
  106. Singer Identity Representation Learning using Self-Supervised Techniques

    Research contributors · arXiv · 2024 · preprint

    Investigates representations for singer identity and whether speech-trained speaker-recognition models transfer adequately to singing voice.

    Cited in Similarity, Memorisation, and Attribution
  107. Sound Recordings Code Appendix on Artificial Intelligence

    SAG-AFTRA and signatory record companies · SAG-AFTRA · 2024 · collective-bargaining-agreement

    Contractual provisions for covered sound recordings, including definitions, consent requirements, compensation, and review of digital voice replicas. Scope is specific to the agreement.

    Cited in Where AI Music Models Are Heading
  108. SoundLabs and Universal Music Group Announce Strategic Agreement to Offer Responsibly Trained AI Technology and Vocal Modeling Plug-In MicDrop to UMG Artists

    Universal Music Group · Universal Music Group · 2024 · official-industry-announcement

    Official announcement describing artist-specific vocal models trained on artists' own voice data, exclusive creative use, ownership control, and artistic approval.

    Cited in Where AI Music Models Are Heading
  109. SoundStream: An End-to-End Neural Audio Codec

    Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · IEEE/ACM Transactions on Audio, Speech, and Language Processing · 2021 · paper

    Original SoundStream paper describing objective and subjective evaluation of an end-to-end neural codec across speech, music, and general audio.

    Cited in Automated Evaluation
  110. SynthID

    Google DeepMind · Google DeepMind · official-technical-documentation

    Official description of watermarking and detection across AI-generated media, including audio generated through Lyria.

    Cited in Where AI Music Models Are Heading
  111. The Belmont Report

    National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research · United States Department of Health and Human Services · 1979 · government-guidance

    Official ethical framework for research involving human participants, including respect for persons, beneficence, justice, and informed consent.

    Cited in Human Listening Tests
  112. The Research Provisions

    Information Commissioner's Office · Information Commissioner's Office · regulatory-guidance

    Official data-protection guidance concerning personal-data processing for research, legal bases, safeguards, transparency, and participant rights.

    Cited in Human Listening Tests
  113. Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio

    Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez · International Society for Music Information Retrieval · 2024 · paper

    Music-specific study comparing raw-audio similarity metrics and proposing the MiRA workflow for replication assessment.

    Cited in What Makes Generated Music Good?
  114. Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio

    Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez · International Society for Music Information Retrieval · 2024 · paper

    Music-specific study evaluating raw-audio similarity metrics and proposing a model-independent replication-assessment workflow.

    Cited in Automated Evaluation
  115. Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio

    Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez · International Society for Music Information Retrieval · 2024 · paper

    Introduces MiRA, a model-independent music-replication assessment workflow using several raw-audio similarity metrics and controlled replication experiments.

    Cited in Similarity, Memorisation, and Attribution
  116. Traces

    OpenTelemetry contributors · OpenTelemetry · documentation

    Official explanation of end-to-end request traces and spans across services.

    Cited in Turning a Model Into a Product
  117. Triton Inference Server Batchers

    NVIDIA · NVIDIA · documentation

    Official documentation for dynamic batching, queue delays, priorities, timeouts, and throughput-oriented request combination.

    Cited in Turning a Model Into a Product
  118. Two Decades of Speaker Recognition Evaluation at the National Institute of Standards and Technology

    Craig Greenberg et al. · National Institute of Standards and Technology · 2020 · government-paper

    Official overview of systematic speaker-recognition and speaker-verification evaluation.

    Cited in Similarity, Memorisation, and Attribution
  119. Universal Music Group and Udio Announce Udio's First Strategic Agreements for New Licensed AI Music Creation Platform

    Universal Music Group · Universal Music Group · 2025 · official-industry-announcement

    Official announcement describing a planned creation platform trained on authorized and licensed music, with fingerprinting, filtering, and protected platform controls.

    Cited in Where AI Music Models Are Heading
  120. Using Dead-Letter Queues in Amazon SQS

    AWS · Amazon Web Services · documentation

    Official guidance for isolating messages that repeatedly fail processing.

    Cited in Turning a Model Into a Product
  121. YuE: Scaling Open Foundation Models for Long-Form Music Generation

    Ruibin Yuan et al. · arXiv · 2025 · paper

    Primary research on long-form lyrics-to-song generation using track-decoupled prediction and structural progressive conditioning.

    Cited in Where AI Music Models Are Heading
Chapter 4: Training and Fine-Tuning a Music Model (113)
  1. A Gentle Introduction to torch.autograd

    PyTorch contributors · PyTorch · documentation

    Official explanation of forward propagation, backward propagation, automatic differentiation, computational graphs, and stored parameter gradients.

    Cited in What Happens During Training?
  2. A Gentle Introduction to torch.autograd

    PyTorch contributors · PyTorch · documentation

    Official explanation of trainable parameters, gradient calculation, and optimisation foundations.

    Cited in Training From Scratch Versus Fine-Tuning
  3. Adam: A Method for Stochastic Optimization

    Diederik P. Kingma and Jimmy Ba · International Conference on Learning Representations · 2015 · paper

    Original paper introducing Adam and its adaptive first-order moment estimates.

    Cited in What Happens During Training?
  4. Additional Guidance for the Training Pipeline

    Google researchers and engineers · Google for Developers · documentation

    Official guidance on profiling input pipelines, offline preprocessing, evaluation intervals, checkpoint selection, and experiment management.

    Cited in Debugging Model Training
  5. Asynchronous Saving with Distributed Checkpoint

    PyTorch contributors · PyTorch · documentation

    Official recipe for reducing synchronous training interruption while writing distributed checkpoints.

    Cited in Compute and Infrastructure
  6. Audio Conditioning for Music Generation via Discrete Bottleneck Features

    Simon Rouard et al. · arXiv · 2024 · paper

    Explores textual inversion using a pretrained text-to-music model and a jointly trained style-conditioning music model with quantised audio features.

    Cited in Training From Scratch Versus Fine-Tuning
  7. Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning

    Fang-Duo Tsai et al. · International Society for Music Information Retrieval · 2024 · paper

    Introduces attention-based audio adapters for a pretrained AudioLDM2 model, with 22 million trainable parameters for timbre transfer, genre transfer, and accompaniment generation.

    Cited in Training From Scratch Versus Fine-Tuning
  8. AudioCraft API Documentation

    Meta AI · Meta AudioCraft · documentation

    Official API documentation for AudioCraft training and inference components including MusicGen, AudioGen, and EnCodec.

    Cited in Compute and Infrastructure
  9. AudioCraft Conditioning Documentation

    Meta AI · Meta AudioCraft · documentation

    Official documentation for text, waveform, joint-embedding, classifier-free guidance, and attribute-dropout conditioning.

    Cited in Debugging Model Training
  10. AudioCraft Language Model API

    Meta AI · Meta AudioCraft · documentation

    Official API for MusicGen language modelling, sampling parameters, classifier-free guidance, condition dropout, and codebook generation.

    Cited in Debugging Model Training
  11. AudioCraft MusicGen Model API

    Meta AI · Meta AudioCraft · documentation

    Official model interface for inference, text conditioning, melody conditioning, continuation, and generation parameters.

    Cited in Training From Scratch Versus Fine-Tuning
  12. AudioCraft Training Documentation

    Meta AI · Meta AudioCraft · documentation

    Official documentation stating that AudioCraft provides training and inference code for MusicGen, AudioGen, and EnCodec.

    Cited in Training From Scratch Versus Fine-Tuning
  13. AudioCraft Training Pipelines

    Meta AI · Meta AudioCraft · documentation

    Official documentation for AudioCraft's PyTorch training pipelines, Flashy, Dora experiment management, environment configuration, grids, clusters, and experiment workflows.

    Cited in Compute and Infrastructure
  14. Autograd Mechanics

    PyTorch contributors · PyTorch · documentation

    Official technical documentation for reverse automatic differentiation and gradient computation.

    Cited in What Happens During Training?
  15. Automatic Mixed Precision

    PyTorch contributors · PyTorch · documentation

    Official recipe for autocasting and gradient scaling in mixed-precision training.

    Cited in Compute and Infrastructure
  16. Automatic Mixed Precision

    PyTorch contributors · PyTorch · documentation

    Official explanation of autocasting, gradient scaling, non-finite gradient handling, and mixed-precision optimiser steps.

    Cited in Debugging Model Training
  17. Automatic Mixed Precision Examples

    PyTorch contributors · PyTorch · documentation

    Official examples covering autocasting, gradient scaling, gradient accumulation, clipping, and mixed-precision training.

    Cited in What Happens During Training?
  18. Automatic Mixed Precision Examples

    PyTorch contributors · PyTorch · documentation

    Official examples for mixed precision, gradient accumulation, unscaling, clipping, and multiple models or losses.

    Cited in Compute and Infrastructure
  19. Automatic Mixed Precision Examples

    PyTorch contributors · PyTorch · documentation

    Official examples for gradient accumulation, unscaling, clipping, and mixed-precision training.

    Cited in Debugging Model Training
  20. Catastrophic Forgetting in the Context of Model Updates

    Rich Harang and Hillary Sanders · arXiv · 2023 · paper

    Compares retraining, fine-tuning, rehearsal, and regularisation approaches for updating models while retaining earlier performance.

    Cited in Training From Scratch Versus Fine-Tuning
  21. Common Pitfalls and Recommended Practices

    scikit-learn contributors · scikit-learn · documentation

    Official guidance on data leakage, inconsistent preprocessing, test separation, and controlling randomness.

    Cited in Overfitting and Memorisation
  22. Common Pitfalls and Recommended Practices

    scikit-learn contributors · scikit-learn · documentation

    Official guidance on data leakage, inconsistent preprocessing, split discipline, and random-state handling.

    Cited in Debugging Model Training
  23. Cross-Validation: Evaluating Estimator Performance

    scikit-learn contributors · scikit-learn · documentation

    Official guidance distinguishing training, validation, and test data and warning that repeated hyperparameter decisions can leak information from the test set.

    Cited in Designing a Small Experiment
  24. Cross-Validation: Evaluating Estimator Performance

    scikit-learn contributors · scikit-learn · documentation

    Official documentation distinguishing training, validation, and test use and explaining how cross-validation estimates generalisation.

    Cited in Overfitting and Memorisation
  25. Cross-Validation: Evaluating Estimator Performance

    scikit-learn contributors · scikit-learn · documentation

    Official guidance for separating training, validation, and test use.

    Cited in Debugging Model Training
  26. CrossEntropyLoss

    PyTorch contributors · PyTorch · documentation

    Official API documentation describing cross-entropy between input logits and targets.

    Cited in What Happens During Training?
  27. CUDA Semantics

    PyTorch contributors · PyTorch · documentation

    Official documentation for CUDA behaviour, the caching allocator, allocated memory, reserved memory, and memory-management functions.

    Cited in Compute and Infrastructure
  28. Data Loading Optimization in PyTorch

    PyTorch contributors · PyTorch · documentation

    Official tutorial covering data-loader workers, pinned memory, prefetching, persistent workers, and data-pipeline throughput.

    Cited in Compute and Infrastructure
  29. Datasets and DataLoaders

    PyTorch contributors · PyTorch · documentation

    Official documentation for datasets, batching, shuffling, sampling, and iterable data loading.

    Cited in What Happens During Training?
  30. Debugging Hangs with Flight Recorder

    PyTorch contributors · PyTorch · documentation

    Official tutorial for diagnosing distributed hangs with stack traces and collective-operation flight-recorder data.

    Cited in Debugging Model Training
  31. Decoupled Weight Decay Regularization

    Ilya Loshchilov and Frank Hutter · International Conference on Learning Representations · 2019 · paper

    Original paper introducing decoupled weight decay for adaptive optimisers, commonly known as AdamW.

    Cited in What Happens During Training?
  32. Deduplicating Training Data Makes Language Models Better

    Katherine Lee et al. · Association for Computational Linguistics · 2022 · paper

    Peer-reviewed language-model study on exact and near duplicate removal, memorised output, train-test overlap, and model efficiency. Used as cross-domain dataset evidence.

    Cited in Overfitting and Memorisation
  33. Deduplicating Training Data Mitigates Privacy Risks in Language Models

    Nikhil Kandpal, Eric Wallace, and Colin Raffel · Proceedings of Machine Learning Research · 2022 · paper

    Peer-reviewed language-model study showing a strong relationship between duplicate count and regeneration frequency. Used here as cross-domain evidence, not as a music-specific result.

    Cited in Overfitting and Memorisation
  34. Deep Learning Tuning Playbook

    Google researchers and engineers · Google for Developers · documentation

    Official practical guide to deep-learning experimentation, tuning, pipeline implementation, and systematic performance improvement.

    Cited in Debugging Model Training
  35. Deep Learning Tuning Playbook FAQ

    Google researchers and engineers · Google for Developers · documentation

    Official guidance on identifying optimisation instability using learning-rate behaviour and gradient norms, with possible responses including warm-up and clipping.

    Cited in Debugging Model Training
  36. DistributedDataParallel

    PyTorch contributors · PyTorch · documentation

    Official API documentation for model replication and gradient synchronisation across data-parallel processes.

    Cited in Compute and Infrastructure
  37. Dropout: A Simple Way to Prevent Neural Networks from Overfitting

    Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · Journal of Machine Learning Research · 2014 · paper

    Original peer-reviewed paper introducing dropout as a neural-network regularisation method.

    Cited in Overfitting and Memorisation
  38. EnCodec Documentation

    Meta AI · Meta AudioCraft · documentation

    Official AudioCraft documentation for training and evaluating the neural audio codec used by MusicGen.

    Cited in Compute and Infrastructure
  39. Exploring Adapter Design Tradeoffs for Low Resource Music Generation

    Atharva Mehta, Shivam Chauhan, and Monojit Choudhury · arXiv · 2025 · paper

    Music-generation study comparing adapter architectures and scales for adapting MusicGen and Mustango to underrepresented music traditions.

    Cited in Training From Scratch Versus Fine-Tuning
  40. Extracting Training Data from Diffusion Models

    Nicholas Carlini et al. · USENIX Security Symposium · 2023 · paper

    Peer-reviewed image-diffusion study demonstrating training-data extraction through generation and filtering. Included as cross-domain evidence about extraction, not as direct evidence about music models.

    Cited in Overfitting and Memorisation
  41. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · NeurIPS · 2022 · paper

    Introduces tiled exact attention that reduces memory reads and writes between GPU memory levels.

    Cited in Compute and Infrastructure
  42. Generation or Replication: Auscultating Audio Latent Diffusion Models

    Dimitrios Bralios et al. · IEEE ICASSP · 2024 · paper

    Peer-reviewed audio study examining memorisation across training-set sizes, duplicated AudioCaps clips, and retrieval metrics for detecting replicated audio.

    Cited in Overfitting and Memorisation
  43. Getting Started with DeviceMesh

    PyTorch contributors · PyTorch · documentation

    Official documentation for organising process groups and multi-dimensional distributed parallelism.

    Cited in Compute and Infrastructure
  44. Getting Started with Distributed Checkpoint

    PyTorch contributors · PyTorch · documentation

    Official tutorial for saving and restoring state partitioned across distributed workers and changing trainer counts.

    Cited in Compute and Infrastructure
  45. Getting Started with Distributed Data Parallel

    PyTorch contributors · PyTorch · documentation

    Official tutorial for multi-process and multi-machine data-parallel training.

    Cited in Compute and Infrastructure
  46. Getting Started with Fully Sharded Data Parallel

    PyTorch contributors · PyTorch · documentation

    Official documentation describing the memory burden of model parameters, gradients, and optimiser states and how these can be sharded for large-scale training.

    Cited in Training From Scratch Versus Fine-Tuning
  47. Getting Started with Fully Sharded Data Parallel

    PyTorch contributors · PyTorch · documentation

    Official tutorial explaining sharding of parameters, gradients, and optimiser states across workers.

    Cited in Compute and Infrastructure
  48. GPU Performance Background User's Guide

    NVIDIA · NVIDIA · documentation

    Official explanation of GPU execution structure, memory hierarchy, and common performance limitations.

    Cited in Compute and Infrastructure
  49. GroupKFold

    scikit-learn contributors · scikit-learn · documentation

    Official cross-validation method for domain-specific groups that should not overlap across folds.

    Cited in Designing a Small Experiment
  50. GroupShuffleSplit

    scikit-learn contributors · scikit-learn · documentation

    Official grouped splitting tool that assigns unique groups rather than individual samples across partitions.

    Cited in Designing a Small Experiment
  51. How to Save Memory by Fusing the Optimizer Step into the Backward Pass

    PyTorch contributors · PyTorch · documentation

    Official tutorial illustrating memory used by parameters, activations, gradients, optimiser states, and optimiser intermediates during training.

    Cited in Compute and Infrastructure
  52. Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

    Or Tal et al. · International Society for Music Information Retrieval · 2024 · paper

    Introduces JASCO for global text conditioning together with local symbolic and audio controls using a new flow-matching architecture and conditioning method.

    Cited in Training From Scratch Versus Fine-Tuning
  53. Large Scale Transformer Model Training with Tensor Parallel

    PyTorch contributors · PyTorch · documentation

    Official tutorial for combining tensor parallelism with fully sharded training for large Transformer models.

    Cited in Compute and Infrastructure
  54. LoRA Conceptual Guide

    Hugging Face contributors · Hugging Face · documentation

    Official conceptual guidance for LoRA matrices, target modules, and parameter-efficient adaptation.

    Cited in Training From Scratch Versus Fine-Tuning
  55. LoRA: Low-Rank Adaptation of Large Language Models

    Edward J. Hu et al. · International Conference on Learning Representations · 2022 · paper

    Original LoRA paper describing frozen pretrained weights with trainable low-rank matrix updates.

    Cited in Training From Scratch Versus Fine-Tuning
  56. MAGNeT Training and Fine-Tuning Documentation

    Meta AI · Meta AudioCraft · documentation

    Official documentation showing checkpoint-based fine-tuning and warning that configurations and changed layers must remain compatible.

    Cited in Training From Scratch Versus Fine-Tuning
  57. Mixed Precision Training

    Paulius Micikevicius et al. · International Conference on Learning Representations · 2018 · paper

    Foundational paper describing lower-precision training, higher-precision weight updates, and loss scaling.

    Cited in Compute and Infrastructure
  58. Mosaic: Memory Profiling for PyTorch

    PyTorch contributors · PyTorch · documentation

    Official tutorial for analysing peak GPU memory, allocation categories, activation checkpointing savings, and memory imbalance.

    Cited in Compute and Infrastructure
  59. Music Transformer

    Cheng-Zhi Anna Huang et al. · Google Brain and Magenta · 2018 · paper

    Original paper describing relative attention for long symbolic music sequences, expressive piano generation, motif continuation, and comparisons with recurrent approaches.

    Cited in Designing a Small Experiment
  60. Music Transformer: Generating Music with Long-Term Structure

    Google Magenta · Google Magenta · 2018 · website

    Official project page explaining the event representation and the contrast between recurrent hidden-state compression and Transformer access to earlier events.

    Cited in Designing a Small Experiment
  61. MusicGen Documentation

    Meta AI · Meta AudioCraft · documentation

    Official overview of the MusicGen model, codec-token representation, and AudioCraft implementation.

    Cited in What Happens During Training?
  62. MusicGen Documentation

    Meta AI · Meta AudioCraft · documentation

    Official overview of MusicGen architecture, pretrained checkpoints, generation modes, and codec-token configuration.

    Cited in Training From Scratch Versus Fine-Tuning
  63. MusicGen Documentation

    Meta AI · Meta AudioCraft · documentation

    Official documentation describing MusicGen's 32 kHz EnCodec tokenizer, four codebooks at 50 Hz, generation modes, and model configuration.

    Cited in Compute and Infrastructure
  64. MusicGen Documentation

    Meta AI · Meta AudioCraft · documentation

    Official documentation for MusicGen architecture, EnCodec token streams, conditioning, and generation.

    Cited in Debugging Model Training
  65. MusicGenSolver API Documentation

    Meta AI · Meta AudioCraft · documentation

    Official implementation showing MusicGen batch preparation, tokenisation, masking, per-codebook cross-entropy, backpropagation, gradient scaling, clipping, optimiser updates, scheduling, and metric logging.

    Cited in What Happens During Training?
  66. MusicGenSolver API Documentation

    Meta AI · Meta AudioCraft · documentation

    Official implementation documentation for MusicGen training, validation, generation, model summaries, metrics, and optimisation.

    Cited in Compute and Infrastructure
  67. MusicGenSolver API Documentation

    Meta AI · Meta AudioCraft · documentation

    Official implementation showing metadata and audio checks, condition preparation, codec encoding, padding masks, per-codebook loss, backward propagation, gradient scaling, clipping, optimiser updates, finite-loss checks, generation, and audio decoding.

    Cited in Debugging Model Training
  68. NVIDIA Deep Learning Performance Guide

    NVIDIA · NVIDIA · documentation

    Official guidance on GPU execution, batch sizing, memory-limited operations, and deep-learning performance.

    Cited in Compute and Infrastructure
  69. Optimizing Model Parameters

    PyTorch contributors · PyTorch · documentation

    Official tutorial covering loss functions, optimisers, zeroing gradients, backward propagation, optimiser steps, and training and validation loops.

    Cited in What Happens During Training?
  70. Overcoming Catastrophic Forgetting in Neural Networks

    James Kirkpatrick et al. · Proceedings of the National Academy of Sciences · 2017 · paper

    Foundational study of catastrophic forgetting during sequential neural-network training and a regularisation method for protecting important earlier parameters.

    Cited in Training From Scratch Versus Fine-Tuning
  71. Parameter-Efficient Fine-Tuning

    Hugging Face contributors · Hugging Face · documentation

    Official PEFT documentation covering efficient adaptation without updating all base-model parameters.

    Cited in Training From Scratch Versus Fine-Tuning
  72. Parameter-Efficient Transfer Learning for Music Foundation Models

    Yiwei Ding and Alexander Lerch · arXiv · 2024 · paper

    Music-specific comparison of adapter-based, prompt-based, and reparameterisation-based transfer methods against probing, fine-tuning, and small models trained from scratch.

    Cited in Training From Scratch Versus Fine-Tuning
  73. Parameter-Efficient Transfer Learning for NLP

    Neil Houlsby et al. · International Conference on Machine Learning · 2019 · paper

    Original paper introducing compact task-specific adapter modules while keeping the original network fixed.

    Cited in Training From Scratch Versus Fine-Tuning
  74. Performance RNN: Generating Music with Expressive Timing and Dynamics

    Ian Simon and Sageev Oore · Google Magenta · 2017 · website

    Official Magenta explanation of an LSTM-based recurrent model using event-based MIDI to generate polyphonic piano performances with expressive timing and dynamics.

    Cited in Designing a Small Experiment
  75. Performance Tuning Guide

    PyTorch contributors · PyTorch · documentation

    Official guidance for improving performance across data loading, memory, precision, and model execution.

    Cited in Compute and Infrastructure
  76. Performance Tuning Guide

    PyTorch contributors · PyTorch · documentation

    Official performance guidance noting that anomaly detection, profilers, and gradcheck are debugging tools rather than normal training settings.

    Cited in Debugging Model Training
  77. pretty_midi Documentation

    Colin Raffel and contributors · Colin Raffel · documentation

    Official documentation for parsing, inspecting, modifying, synthesising, and extracting information from MIDI files.

    Cited in Designing a Small Experiment
  78. Reproducibility

    PyTorch contributors · PyTorch · documentation

    Official guidance on random seeds, deterministic behaviour, DataLoader worker seeding, and the limits of reproducibility across releases and platforms.

    Cited in Designing a Small Experiment
  79. Reproducibility

    PyTorch contributors · PyTorch · documentation

    Official guidance on seeds, deterministic algorithms, data-loader randomness, and limits to exact reproducibility across releases and platforms.

    Cited in Compute and Infrastructure
  80. Reproducibility

    PyTorch contributors · PyTorch · documentation

    Official guidance on random seeds, deterministic behaviour, data-loader randomness, and reproducibility limitations.

    Cited in Debugging Model Training
  81. Saving and Loading Models

    PyTorch contributors · PyTorch · documentation

    Official guidance for saving model state and general training checkpoints.

    Cited in What Happens During Training?
  82. Saving and Loading Models

    Matthew Inkawhich and PyTorch contributors · PyTorch · documentation

    Official guidance for saving weights, general checkpoints, optimiser state, and restoring models.

    Cited in Compute and Infrastructure
  83. Saving and Loading Models

    PyTorch contributors · PyTorch · documentation

    Official guidance for preserving model, optimiser, scheduler, and training state.

    Cited in Debugging Model Training
  84. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing a Transformer language model trained over several streams of compressed music tokens.

    Cited in What Happens During Training?
  85. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing a music language model trained over compressed EnCodec token streams.

    Cited in Training From Scratch Versus Fine-Tuning
  86. Simple and Controllable Music Generation

    Jade Copet et al. · NeurIPS · 2023 · paper

    Original MusicGen paper. Its memorisation experiment tests exact and 80 percent partial matches in the first codec stream for five-second continuations from 20,000 sampled training examples.

    Cited in Overfitting and Memorisation
  87. Simple and Controllable Music Generation

    Jade Copet et al. · NeurIPS · 2023 · paper

    Original MusicGen paper describing a single-stage Transformer over multiple streams of compressed music tokens.

    Cited in Compute and Infrastructure
  88. Simple and Controllable Music Generation

    Jade Copet et al. · NeurIPS · 2023 · paper

    Original MusicGen paper describing conditional autoregressive modelling over several streams of compressed music tokens.

    Cited in Debugging Model Training
  89. The Lakh MIDI Dataset v0.1

    Colin Raffel · Colin Raffel · dataset

    Official project page describing 176,581 unique MIDI files, including 45,129 matched and aligned with Million Song Dataset entries.

    Cited in Designing a Small Experiment
  90. The MAESTRO Dataset

    Curtis Hawthorne et al. · Google Magenta · dataset

    Official dataset page describing roughly 200 hours of aligned piano audio and MIDI, expressive performance information, metadata, predefined splits, and licence terms.

    Cited in Designing a Small Experiment
  91. torch.autograd

    PyTorch contributors · PyTorch · documentation

    Official automatic-differentiation documentation relevant to gradients and backward-pass state.

    Cited in Compute and Infrastructure
  92. torch.autograd

    PyTorch contributors · PyTorch · documentation

    Official automatic-differentiation documentation, including anomaly detection and gradient-checking tools.

    Cited in Debugging Model Training
  93. torch.isfinite

    PyTorch contributors · PyTorch · documentation

    Official API for identifying finite and non-finite tensor elements.

    Cited in Debugging Model Training
  94. torch.nn.Dropout

    PyTorch contributors · PyTorch · documentation

    Official API documentation explaining that dropout randomly zeros activations during training and behaves differently during evaluation.

    Cited in Overfitting and Memorisation
  95. torch.nn.Module

    PyTorch contributors · PyTorch · documentation

    Official module API containing forward and backward hook mechanisms for inspecting model computations.

    Cited in Debugging Model Training
  96. torch.nn.utils.clip_grad_norm_

    PyTorch contributors · PyTorch · documentation

    Official documentation for calculating and clipping the total norm of parameter gradients.

    Cited in What Happens During Training?
  97. torch.nn.utils.clip_grad_norm_

    PyTorch contributors · PyTorch · documentation

    Official API for calculating and clipping the total gradient norm.

    Cited in Debugging Model Training
  98. torch.optim

    PyTorch contributors · PyTorch · documentation

    Official documentation for optimisation algorithms and learning-rate schedulers.

    Cited in What Happens During Training?
  99. torch.profiler

    PyTorch contributors · PyTorch · documentation

    Official profiler documentation for tracing CPU and CUDA operations, shapes, stacks, and memory.

    Cited in Compute and Infrastructure
  100. torch.profiler

    PyTorch contributors · PyTorch · documentation

    Official profiler for collecting performance and memory information during training and inference.

    Cited in Debugging Model Training
  101. torch.Tensor.register_hook

    PyTorch contributors · PyTorch · documentation

    Official API for registering callbacks executed when tensor gradients are calculated.

    Cited in Debugging Model Training
  102. torch.utils.checkpoint

    PyTorch contributors · PyTorch · documentation

    Official documentation explaining activation checkpointing as a compute-for-memory trade-off.

    Cited in Compute and Infrastructure
  103. torch.utils.data

    PyTorch contributors · PyTorch · documentation

    Official dataset, sampler, batching, worker, prefetch, and persistent-worker documentation.

    Cited in Compute and Infrastructure
  104. torch.utils.data

    PyTorch contributors · PyTorch · documentation

    Official documentation for datasets, batching, sampling, workers, prefetching, and data loading.

    Cited in Debugging Model Training
  105. Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio

    Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez · International Society for Music Information Retrieval · 2024 · paper

    Peer-reviewed music-specific work proposing the MiRA assessment approach and evaluating several raw-audio similarity metrics in controlled replication tests.

    Cited in Overfitting and Memorisation
  106. Transfer Learning for Computer Vision Tutorial

    Sasank Chilamkurthy and PyTorch contributors · PyTorch · documentation

    Official PyTorch tutorial contrasting full fine-tuning with freezing a pretrained network and training only a final layer.

    Cited in Training From Scratch Versus Fine-Tuning
  107. Underfitting Versus Overfitting

    scikit-learn contributors · scikit-learn · documentation

    Official example illustrating underfitting, suitable complexity, overfitting, and cross-validated error.

    Cited in Overfitting and Memorisation
  108. Understanding CUDA Memory Usage

    PyTorch contributors · PyTorch · documentation

    Official guidance for recording and visualising CUDA memory snapshots and investigating allocation patterns.

    Cited in Compute and Infrastructure
  109. Validation Curves: Plotting Scores to Evaluate Models

    scikit-learn contributors · scikit-learn · documentation

    Official documentation explaining training and validation curves, underfitting, overfitting, and how additional data can affect generalisation.

    Cited in Overfitting and Memorisation
  110. Visualizing Models, Data, and Training with TensorBoard

    PyTorch contributors · PyTorch · documentation

    Official tutorial for recording and visualising training metrics, images, model graphs, and embeddings.

    Cited in Compute and Infrastructure
  111. What Is torch.nn Really?

    PyTorch contributors · PyTorch · documentation

    Official tutorial distinguishing training and evaluation modes and calculating validation loss.

    Cited in What Happens During Training?
  112. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · SC20 · 2020 · paper

    Foundational work on eliminating replicated parameter, gradient, and optimiser-state memory in distributed training.

    Cited in Compute and Infrastructure
  113. Zeroing Out Gradients in PyTorch

    PyTorch contributors · PyTorch · documentation

    Official explanation of gradient accumulation and the need to clear gradients between ordinary training steps.

    Cited in What Happens During Training?
Chapter 2: Building the Musical Training Dataset (51)
  1. About CC Licenses

    Creative Commons · Creative Commons · documentation

    Official overview of Creative Commons licence elements and the permissions and conditions attached to standard CC licences.

    Cited in Licensing, Consent, and Provenance
  2. Announcing AudioSet: A Dataset for Audio Event Research

    Google Research · Google Research · 2017 · website

    Official announcement describing AudioSet as more than two million ten-second excerpts labelled with hundreds of sound-event categories.

    Cited in What Is a Training Example?
  3. Audio Based Disambiguation of Music Genre Tags

    Romain Hennequin, Jimena Royo-Letelier, and Manuel Moussallam · arXiv · 2018 · paper

    Research showing how audio-based genre embeddings can identify duplicate labels, translate between tag systems, and recover genre relationships.

    Cited in Metadata and Captioning
  4. Audio Resampling

    Caroline Chen and Moto Hira · PyTorch · documentation

    Official tutorial on bandlimited audio resampling, filter parameters, roll-off, aliasing, and efficient reuse of resampling kernels.

    Cited in Collecting and Cleaning Audio
  5. Audio Set: An Ontology and Human-Labeled Dataset for Audio Events

    Jort F. Gemmeke et al. · Google Research · 2017 · paper

    Original paper describing a large-scale dataset of manually labelled audio events organised through an ontology.

    Cited in What Is a Training Example?
  6. Audio Set: An Ontology and Human-Labeled Dataset for Audio Events

    Jort F. Gemmeke et al. · Google Research · 2017 · paper

    Original paper describing a manually annotated, multi-label audio-event dataset organised through a structured ontology.

    Cited in Metadata and Captioning
  7. Audio Set: An Ontology and Human-Labeled Dataset for Audio Events

    Jort F. Gemmeke et al. · Google Research · 2017 · paper

    Original paper introducing AudioSet, its sound-event ontology, and its large multi-label collection with highly uneven class frequencies.

    Cited in Dataset Balance and Bias
  8. Bias beyond Borders: Global Inequalities in AI-Generated Music

    Ahmet Solak, Florian Grötschla, Luca A. Lanzendörfer, and Roger Wattenhofer · arXiv · 2025 · paper

    Introduces the globally balanced GlobalDISCO evaluation dataset and reports disparities across regions, languages, and mainstream versus geographically niche genres.

    Cited in Dataset Balance and Bias
  9. Can MusicGen Create Training Data for MIR Tasks?

    Nadine Kroher, Helena Cuesta, and Aggelos Pikrakis · arXiv · 2023 · paper

    Early experiment using MusicGen to create genre-conditioned synthetic training data for a music information retrieval classifier.

    Cited in Dataset Balance and Bias
  10. Chromaprint

    Lukas Lalinsky and contributors · AcoustID · documentation

    Official repository describing Chromaprint as a compact audio fingerprint system designed for near-identical full-file identification and duplicate audio detection.

    Cited in Collecting and Cleaning Audio
  11. Copyright and Artificial Intelligence, Part 3: Generative AI Training

    U.S. Copyright Office · U.S. Copyright Office · 2025 · documentation

    Official report describing the legal and policy debate around copyrighted works in generative AI training and the fact-specific, jurisdiction-sensitive nature of the analysis.

    Cited in Licensing, Consent, and Provenance
  12. Copyright Registration of Musical Compositions and Sound Recordings

    U.S. Copyright Office · U.S. Copyright Office · 2020 · documentation

    Official circular explaining that musical compositions and sound recordings are separate works and identifying the different material and authors associated with each.

    Cited in Licensing, Consent, and Provenance
  13. Creative Commons Frequently Asked Questions

    Creative Commons · Creative Commons · 2025 · documentation

    Official FAQ explaining that CC licences do not automatically clear third-party publicity, privacy, personality, trademark, patent, or other rights outside their scope.

    Cited in Licensing, Consent, and Provenance
  14. Croissant Format Specification 1.1

    MLCommons Croissant Working Group · MLCommons · 2026 · documentation

    Machine-learning dataset metadata standard supporting licence fields and PROV-O relationships at dataset, resource, record, field, and value level.

    Cited in Licensing, Consent, and Provenance
  15. DALI: A Large Dataset of Synchronized Audio, Lyrics and Notes

    Gabriel Meseguer-Brocal, Alice Cohen-Hadria, and Geoffroy Peeters · arXiv · 2019 · paper

    Original paper introducing 5,358 audio tracks with time-aligned lyrics and vocal melody notes at four levels of granularity.

    Cited in What Is a Training Example?
  16. Data Leakage in Cross-Modal Retrieval Training: A Case Study

    Benno Weck and Xavier Serra · Music Technology Group, Universitat Pompeu Fabra · 2023 · paper

    Audio-specific study finding duplicates across SoundDesc train and evaluation data, showing overly optimistic retrieval performance, and proposing corrected splits.

    Cited in Collecting and Cleaning Audio
  17. Dataset Balancing Can Hurt Model Performance

    R. Channing Moore, Daniel P. W. Ellis, Eduardo Fonseca, Shawn Hershey, Aren Jansen, and Manoj Plakal · arXiv · 2023 · paper

    Audio-specific study showing that balancing AudioSet improved one public evaluation score while hurting a separate evaluation set and did not clearly improve rare-class performance.

    Cited in Dataset Balance and Bias
  18. Datasheets for Datasets

    Timnit Gebru et al. · Communications of the ACM · 2021 · paper

    Foundational framework proposing documentation of dataset motivation, composition, collection, processing, uses, distribution, and maintenance.

    Cited in Dataset Balance and Bias
  19. Datasheets for Datasets

    Timnit Gebru et al. · Communications of the ACM · 2021 · paper

    Foundational dataset-documentation framework covering motivation, composition, collection, processing, distribution, maintenance, uses, and limitations.

    Cited in Licensing, Consent, and Provenance
  20. Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset

    Curtis Hawthorne et al. · Google Magenta · 2018 · paper

    Original paper introducing paired piano audio and MIDI with fine alignment for transcription, symbolic composition, and waveform synthesis.

    Cited in What Is a Training Example?
  21. Ethical Dimensions of Music Information Retrieval Technology

    Andre Holzapfel et al. · Transactions of the International Society for Music Information Retrieval · 2018 · paper

    Music-specific discussion of cultural, social, data, and algorithmic biases in music information retrieval research and technology.

    Cited in Dataset Balance and Bias
  22. ffprobe Documentation

    FFmpeg Project · FFmpeg · documentation

    Official documentation for inspecting media containers and streams, selecting audio streams, checking usable stream configuration, and producing structured output.

    Cited in Collecting and Cleaning Audio
  23. Jukebox: A Generative Model for Music

    Prafulla Dhariwal et al. · OpenAI · 2020 · paper

    Original Jukebox paper describing music generation conditioned on artist, genre, and unaligned lyrics.

    Cited in What Is a Training Example?
  24. Leveraging Knowledge Bases and Parallel Annotations for Music Genre Translation

    Elena V. Epure, Anis Khlif, and Romain Hennequin · arXiv · 2019 · paper

    Research addressing diversity and subjectivity across genre tag systems and methods for translating annotations between taxonomies.

    Cited in Metadata and Captioning
  25. librosa.effects.split

    librosa development team · librosa · documentation

    Official documentation for identifying non-silent intervals in mono or multichannel audio.

    Cited in Collecting and Cleaning Audio
  26. librosa.effects.trim

    librosa development team · librosa · documentation

    Official documentation for trimming leading and trailing silence using a threshold below a reference level.

    Cited in Collecting and Cleaning Audio
  27. librosa.resample

    librosa development team · librosa · documentation

    Official documentation for converting a time series from an original sample rate to a target sample rate.

    Cited in Collecting and Cleaning Audio
  28. LP-MusicCaps

    SeungHeon Doh and contributors · KAIST Music and Audio Computing Lab · 2023 · documentation

    Official project repository documenting tag-to-caption generation and audio-to-caption training.

    Cited in Metadata and Captioning
  29. LP-MusicCaps: LLM-Based Pseudo Music Captioning

    SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam · International Society for Music Information Retrieval · 2023 · paper

    Original paper generating large-scale captions from music tags and evaluating caption quality through correct attributes and incorrect attributes.

    Cited in Metadata and Captioning
  30. Missing Melodies: AI Music Generation and its Nearly Complete Omission of the Global South

    Atharva Mehta, Shivam Chauhan, and Monojit Choudhury · arXiv · 2024 · paper

    Large review of more than one million dataset hours and over 200 papers examining geographic and cultural representation in music-generation research.

    Cited in Dataset Balance and Bias
  31. Music Classification Annotations

    Music Technology Group · Music Technology Group, Universitat Pompeu Fabra · dataset

    Official annotation documentation describing raw labels from three annotators and a clean subset containing agreed labels while excluding unmatched and instrumental answers where relevant.

    Cited in Metadata and Captioning
  32. Music for All: Exploring Multicultural Representations in Music Generation Models

    Atharva Mehta, Shivam Chauhan, Amirbek Djanibekov, Atharva Kulkarni, Gus Xia, and Monojit Choudhury · arXiv · 2025 · paper

    Music-specific study measuring genre and cultural underrepresentation in generation datasets and testing parameter-efficient adaptation for Hindustani classical and Turkish makam music.

    Cited in Dataset Balance and Bias
  33. MusicCaps Dataset Card

    Google · Google · 2023 · dataset

    Official dataset card describing 5,521 records and fields including source video ID, segment timing, AudioSet labels, aspect lists, and musician-written captions.

    Cited in What Is a Training Example?
  34. MusicCaps Dataset Card

    Google · Google · 2023 · dataset

    Official dataset card describing 5,521 ten-second music examples with AudioSet labels, aspect lists, and free-text captions written by musicians.

    Cited in Metadata and Captioning
  35. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original MusicLM paper introducing text-conditioned music generation and the MusicCaps dataset of about 5.5 thousand music-text pairs.

    Cited in What Is a Training Example?
  36. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original MusicLM paper describing MusicCaps, expert-written captions, aspect lists, text conditioning, and caption adherence evaluation.

    Cited in Metadata and Captioning
  37. Position: Measure Dataset Diversity, Don't Just Claim It

    Sabina Tomkins et al. · arXiv · 2024 · paper

    Argues that claims about dataset diversity need explicit definitions, dimensions, and measurements rather than broad unsupported descriptions.

    Cited in Dataset Balance and Bias
  38. PROV-O: The PROV Ontology

    W3C Provenance Working Group · World Wide Web Consortium · 2013 · documentation

    W3C Recommendation providing classes and relationships for representing and exchanging provenance information across systems.

    Cited in Licensing, Consent, and Provenance
  39. Recommendation ITU-R BS.1770-5: Algorithms to Measure Audio Programme Loudness and True-Peak Audio Level

    ITU Radiocommunication Sector · International Telecommunication Union · 2023 · documentation

    Official recommendation defining objective programme loudness and true-peak measurement algorithms.

    Cited in Collecting and Cleaning Audio
  40. Regulation (EU) 2016/679, General Data Protection Regulation

    European Parliament and Council of the European Union · European Union · 2016 · documentation

    Official EU regulation defining consent and, where processing relies on consent, requiring that it be demonstrable and withdrawable under the stated conditions.

    Cited in Licensing, Consent, and Provenance
  41. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing a music generation model conditioned on text descriptions or melodic features.

    Cited in What Is a Training Example?
  42. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing full-length music standardised to 32 kHz, mono downmixing in the main setup, and random 30-second training crops.

    Cited in Collecting and Cleaning Audio
  43. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing music generation conditioned on textual descriptions or melodic features.

    Cited in Metadata and Captioning
  44. Stable Audio Open

    Zach Evans et al. · Stability AI · 2024 · paper

    Original paper describing an open text-to-audio model producing stereo audio at 44.1 kHz.

    Cited in Collecting and Cleaning Audio
  45. Stable Audio Tools Dataset Documentation

    Stability AI · Stability AI · documentation

    Official dataset documentation describing loading and resampling, optional LUFS normalisation, silence handling, padding and cropping, padding masks, and channel conversion.

    Cited in Collecting and Cleaning Audio
  46. The MAESTRO Dataset

    Google Magenta · Google Magenta · 2018 · dataset

    Official dataset page describing paired audio and MIDI, key velocities, pedal positions, descriptive metadata, and alignment of about 3 milliseconds.

    Cited in What Is a Training Example?
  47. The MAESTRO Dataset

    Google Magenta · Google Magenta · 2018 · dataset

    Official dataset page stating that MAESTRO is available under the Creative Commons Attribution Non-Commercial Share-Alike 4.0 licence.

    Cited in Licensing, Consent, and Provenance
  48. The MTG-Jamendo Dataset

    Dmitry Bogdanov et al. · Music Technology Group, Universitat Pompeu Fabra · 2019 · dataset

    Official dataset page describing more than 55,000 tracks and 195 uploader-provided genre, instrument, and mood or theme tags.

    Cited in Metadata and Captioning
  49. The MTG-Jamendo Dataset

    Dmitry Bogdanov et al. · Music Technology Group, Universitat Pompeu Fabra · 2019 · dataset

    Official dataset page describing more than 55,000 tracks, 195 genre, instrument, and mood or theme tags, uploader-provided labels, published splits, and tag distributions.

    Cited in Dataset Balance and Bias
  50. Using Audio Fingerprinting for Duplicate Detection and Thumbnail Generation

    Christopher J. C. Burges, John C. Platt, and Soumya Jana · Microsoft Research · 2005 · paper

    Research describing audio fingerprinting for detecting duplicate clips even when files differ in compression quality or duration.

    Cited in Collecting and Cleaning Audio
  51. Using CC-Licensed Works for AI Training

    Creative Commons · Creative Commons · 2025 · documentation

    Official guidance explaining that copyright treatment of AI training varies and discussing conservative handling of attribution, share-alike, noncommercial, and no-derivatives licence elements.

    Cited in Licensing, Consent, and Provenance
Chapter 1: Turning Music Into Data (31)
  1. About MIDI Part 3: MIDI Messages

    The MIDI Association · The MIDI Association · documentation

    Official overview of note on, note off, velocity, program change, pitch bend, aftertouch, and control change messages.

    Cited in Three Ways to Represent Music
  2. About MIDI Part 4: MIDI Files

    The MIDI Association · The MIDI Association · documentation

    Official explanation that Standard MIDI Files store timed instructions and events rather than recorded sound, with playback produced by a connected sound engine.

    Cited in Three Ways to Represent Music
  3. Audio Set: An Ontology and Human-Labeled Dataset for Audio Events

    Jort F. Gemmeke et al. · Google Research · 2017 · paper

    Original paper introducing the AudioSet ontology and large-scale human-labelled dataset used for audio event recognition.

    Cited in Audio Embeddings
  4. AudioLM: A Language Modeling Approach to Audio Generation

    Zalan Borsos et al. · Google Research · 2022 · paper

    Original AudioLM paper describing audio generation as language modelling over discrete audio tokens and combining semantic and acoustic token types.

    Cited in Audio Tokens and Neural Codecs
  5. CLAP: Learning Audio Concepts From Natural Language Supervision

    Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · Microsoft Research · 2022 · paper

    Original CLAP paper describing separate audio and text encoders trained contrastively in a shared multimodal embedding space.

    Cited in Audio Embeddings
  6. Core Audio Essentials

    Apple Developer Documentation · 2017 · documentation

    Official overview of audio data formats, including sample rate, bits per channel, and channels per frame.

    Cited in What Is Sound?
  7. Digitizing Audio

    Adobe · 2021 · documentation

    Official documentation covering microphones, digital sampling, sample rate, bit depth, amplitude values, channels, and audio file size.

    Cited in What Is Sound?
  8. Displaying Audio in the Waveform Editor

    Adobe · 2021 · documentation

    Official documentation explaining waveform and spectral displays, their axes, amplitude peaks, frequency components, colour intensity, and the time-frequency resolution tradeoff.

    Cited in Three Ways to Represent Music
  9. Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset

    Curtis Hawthorne et al. · Google Magenta · 2019 · paper

    Original MAESTRO paper presenting aligned audio and note events for transcription, symbolic composition, and waveform synthesis.

    Cited in Three Ways to Represent Music
  10. EnCodec: High Fidelity Neural Audio Compression

    Meta Research · Meta Research · documentation

    Official repository documenting available 24 kHz and 48 kHz EnCodec models and supported bitrate settings.

    Cited in Audio Tokens and Neural Codecs
  11. GANSynth: Adversarial Neural Audio Synthesis

    Jesse Engel et al. · Google Magenta · 2019 · paper

    Original paper describing neural audio synthesis in the spectral domain using log magnitudes and instantaneous frequencies.

    Cited in Three Ways to Represent Music
  12. High Fidelity Neural Audio Compression

    Alexandre Defossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · Meta AI · 2022 · paper

    Original EnCodec paper describing a real-time encoder-decoder codec with a quantized latent space and evaluations across speech, music, mono, and stereo audio.

    Cited in Audio Tokens and Neural Codecs
  13. High-Fidelity Audio Compression with Improved RVQGAN

    Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · Descript · 2023 · paper

    Original Descript Audio Codec paper describing high-fidelity universal audio compression into discrete codes using improved residual vector quantization and reconstruction training.

    Cited in Audio Tokens and Neural Codecs
  14. How Do We Hear?

    National Institute on Deafness and Other Communication Disorders · 2022 · website

    Official explanation of how sound waves produce vibrations that are converted into electrical signals in the auditory system.

    Cited in What Is Sound?
  15. MuLan: A Joint Embedding of Music Audio and Natural Language

    Qingqing Huang et al. · Google Research · 2022 · paper

    Original MuLan paper describing a two-tower joint music and text embedding model for retrieval, tagging, transfer learning, and zero-shot tasks.

    Cited in Audio Embeddings
  16. Music Transformer

    Cheng-Zhi Anna Huang et al. · Google Brain · 2018 · paper

    Original paper describing a Transformer model for long symbolic piano performances, continuations, and melody-conditioned accompaniment.

    Cited in Three Ways to Represent Music
  17. Music Transformer: Generating Music with Long-Term Structure

    Google Magenta · Google Magenta · 2018 · website

    Official project overview and generated examples for Music Transformer.

    Cited in Three Ways to Represent Music
  18. MusicGen: Simple and Controllable Music Generation

    Meta AI · Meta AudioCraft · documentation

    Official implementation documentation describing MusicGen's 32 kHz EnCodec tokenizer and model configuration.

    Cited in What Is Sound?
  19. MusicGen: Simple and Controllable Music Generation

    Meta AI · Meta AudioCraft · documentation

    Official implementation documentation describing MusicGen's 32 kHz EnCodec tokenizer, four codebooks, and 50 Hz token rate.

    Cited in Audio Tokens and Neural Codecs
  20. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original MusicLM paper explaining how MuLan music embeddings were used during training and MuLan text embeddings were used as conditioning during generation.

    Cited in Audio Embeddings
  21. scipy.signal.spectrogram

    SciPy · documentation

    Official SciPy documentation describing a spectrogram as a view of changing frequency content over time and explaining its segmented Fourier transform parameters.

    Cited in Three Ways to Represent Music
  22. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing compressed discrete music representations, common music sample rates, and mono and stereo generation.

    Cited in What Is Sound?
  23. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing one Transformer language model operating over several streams of compressed discrete music tokens.

    Cited in Audio Tokens and Neural Codecs
  24. Sound Classification with YAMNet

    TensorFlow · 2024 · documentation

    Official TensorFlow Hub tutorial describing YAMNet, its AudioSet classes, audio input requirements, and its score, embedding, and spectrogram outputs.

    Cited in Audio Embeddings
  25. SoundStream: An End-to-End Neural Audio Codec

    Neil Zeghidour et al. · Google Research · 2021 · paper

    Original paper describing an end-to-end encoder, residual vector quantizer, and decoder for speech, music, and general audio, including variable bitrate operation.

    Cited in Audio Tokens and Neural Codecs
  26. The MAESTRO Dataset

    Google Magenta · Google Magenta · 2018 · dataset

    Official dataset page describing about 200 hours of paired piano audio and MIDI aligned to about 3 milliseconds, including velocity and pedal data.

    Cited in Three Ways to Represent Music
  27. The Science of Sound

    NASA · website

    Official overview of sound as a longitudinal wave and the relationship of amplitude and frequency to loudness and pitch.

    Cited in What Is Sound?
  28. Transfer Learning with YAMNet for Environmental Sound Classification

    TensorFlow · 2024 · documentation

    Official TensorFlow tutorial describing YAMNet's 1,024-dimensional embeddings and their use as pretrained features for a smaller audio classifier.

    Cited in Audio Embeddings
  29. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

    Leland McInnes, John Healy, and James Melville · arXiv · 2018 · paper

    Original UMAP paper describing a scalable dimensionality reduction method used to project high-dimensional data for visualisation.

    Cited in Audio Embeddings
  30. WAVEFORMATEX Structure

    Microsoft Learn · 2021 · documentation

    Official format documentation defining channel count, samples per second, and bits per sample for waveform audio.

    Cited in What Is Sound?
  31. WaveNet: A Generative Model for Raw Audio

    Aaron van den Oord et al. · Google DeepMind · 2016 · paper

    Original paper introducing an autoregressive model that generates raw audio waveforms one sample at a time and includes experiments on music.

    Cited in Three Ways to Represent Music
Chapter 3: How an AI Model Generates Music (59)
  1. Audio Conditioning for Music Generation via Discrete Bottleneck Features

    Simon Rouard, Yossi Adi, Jade Copet, Axel Roebel, and Alexandre Défossez · arXiv · 2024 · paper

    Introduces audio-based style conditioning, joint text and audio control, and separate guidance over the two condition types.

    Cited in Conditioning and Control
  2. AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

    Haohe Liu et al. · ICML · 2023 · paper

    Introduces text-to-audio generation through latent diffusion, an audio autoencoder, and aligned CLAP audio and text embeddings.

    Cited in Diffusion Models
  3. AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

    Haohe Liu et al. · ICML · 2023 · paper

    Describes text-guided latent audio diffusion using aligned audio and text embeddings.

    Cited in Conditioning and Control
  4. AudioLM Examples

    Google Research · Google Research · 2022 · website

    Official project page with speech and piano continuation examples and the AudioLM abstract.

    Cited in Prediction as the Basic Learning Task
  5. AudioLM Examples

    Google Research · Google Research · 2022 · website

    Official project page containing speech and piano audio-continuation examples.

    Cited in Hierarchical Generation
  6. AudioLM: A Language Modeling Approach to Audio Generation

    Zalan Borsos et al. · Google Research · 2022 · paper

    Original paper mapping audio to discrete tokens and treating audio generation as language modelling, including piano continuations from audio prompts.

    Cited in Prediction as the Basic Learning Task
  7. AudioLM: A Language Modeling Approach to Audio Generation

    Zalan Borsos et al. · Google Research · 2022 · paper

    Original paper describing hierarchical autoregressive language models over semantic and neural codec tokens for long-term audio consistency and acoustic quality.

    Cited in Autoregressive Music Models
  8. AudioLM: A Language Modeling Approach to Audio Generation

    Zalan Borsos et al. · Google Research · 2022 · paper

    Original paper describing three-stage hierarchical generation with semantic, coarse acoustic, and fine acoustic tokens.

    Cited in Hierarchical Generation
  9. Classifier-Free Diffusion Guidance

    Jonathan Ho and Tim Salimans · arXiv · 2022 · paper

    Introduces guidance based on combined conditional and unconditional diffusion predictions without requiring a separate classifier.

    Cited in Diffusion Models
  10. Classifier-Free Diffusion Guidance

    Jonathan Ho and Tim Salimans · arXiv · 2022 · paper

    Introduces guidance based on combined conditional and unconditional predictions.

    Cited in Conditioning and Control
  11. CrossEntropyLoss

    PyTorch contributors · PyTorch · documentation

    Official documentation for computing cross-entropy loss between model logits and target classes.

    Cited in Prediction as the Basic Learning Task
  12. Denoising Diffusion Implicit Models

    Jiaming Song, Chenlin Meng, and Stefano Ermon · ICLR · 2021 · paper

    Introduces non-Markovian diffusion sampling processes that use the same training objective while supporting faster generation with fewer sampling steps.

    Cited in Diffusion Models
  13. Denoising Diffusion Probabilistic Models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel · NeurIPS · 2020 · paper

    Foundational paper describing a Gaussian forward corruption process, learned reverse denoising process, and noise-prediction training objective.

    Cited in Diffusion Models
  14. DiffWave: A Versatile Diffusion Model for Audio Synthesis

    Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · ICLR · 2021 · paper

    Original paper describing conditional and unconditional waveform generation by converting white noise into structured audio through diffusion.

    Cited in Diffusion Models
  15. Fast Timing-Conditioned Latent Audio Diffusion

    Zach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley, and Jordi Pons · arXiv · 2024 · paper

    Describes Stable Audio's text and timing conditioning for long-form, variable-length stereo generation.

    Cited in Conditioning and Control
  16. High-Resolution Image Synthesis with Latent Diffusion Models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · CVPR · 2022 · paper

    Foundational latent diffusion paper explaining diffusion in the compressed representation of a pretrained autoencoder to reduce computational requirements.

    Cited in Diffusion Models
  17. Improved Denoising Diffusion Probabilistic Models

    Alexander Quinn Nichol and Prafulla Dhariwal · ICML · 2021 · paper

    Introduces improvements including learned reverse variances and an alternative noise schedule, with substantially fewer sampling passes in tested settings.

    Cited in Diffusion Models
  18. JASCO: Joint Audio and Symbolic Conditioning

    Meta AI · Meta AudioCraft · 2024 · documentation

    Official AudioCraft documentation for temporally controlled generation using text, symbolic, and audio conditions.

    Cited in Conditioning and Control
  19. Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

    Or Tal, Alon Ziv, Itai Gat, Felix Kreuk, and Yossi Adi · arXiv · 2024 · paper

    Introduces JASCO, combining global text with time-aligned chord, melody, drum, and full-mix controls.

    Cited in Conditioning and Control
  20. Jukebox: A Generative Model for Music

    Prafulla Dhariwal et al. · OpenAI · 2020 · paper

    Original paper describing a multi-scale VQ-VAE and autoregressive Transformer priors for generating music with singing over multiple minutes.

    Cited in Autoregressive Music Models
  21. Jukebox: A Generative Model for Music

    Prafulla Dhariwal et al. · OpenAI · 2020 · paper

    Original paper describing multi-scale VQ-VAE compression, a top-level autoregressive prior, and middle and bottom upsampling models.

    Cited in Hierarchical Generation
  22. Jukebox: A Generative Model for Music

    Prafulla Dhariwal et al. · OpenAI · 2020 · paper

    Describes artist, genre, and unaligned lyric conditioning for raw-audio song generation.

    Cited in Conditioning and Control
  23. Jukebox: A Generative Model for Music

    Prafulla Dhariwal et al. · OpenAI · 2020 · paper

    Describes a multi-scale VQ-VAE and autoregressive Transformer priors for raw-audio music generation with reported coherence over multiple minutes.

    Cited in Why Long Songs Are Difficult
  24. Long-form Music Generation with Latent Diffusion

    Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · arXiv · 2024 · paper

    Describes a timing-conditioned diffusion Transformer operating on a highly downsampled continuous audio latent for long-form text-to-music generation.

    Cited in Diffusion Models
  25. Long-form Music Generation with Latent Diffusion

    Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · arXiv · 2024 · paper

    Describes a diffusion Transformer operating at a latent rate of 21.5 positions per second, trained on long temporal contexts and generating music up to 4 minutes and 45 seconds.

    Cited in Why Long Songs Are Difficult
  26. MELONS: Generating Melody with Long-Term Structure Using Transformers and Structure Graph

    Yi Zou et al. · ICASSP · 2022 · paper

    Introduces a graph representation of bar-level melody relationships and separates structure generation from structure-conditioned melody generation.

    Cited in Why Long Songs Are Difficult
  27. Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion

    Flavio Schneider, Ojasv Kamal, Zhijing Jin, and Bernhard Schölkopf · arXiv · 2023 · paper

    Describes a two-stage system using a diffusion autoencoder and text-conditioned latent diffusion for long stereo music generation.

    Cited in Diffusion Models
  28. MuseCoco: Generating Symbolic Music from Text

    Peiling Lu et al. · arXiv · 2023 · paper

    Introduces text-to-symbolic-music generation using explicit musical attributes as an intermediate control representation.

    Cited in Conditioning and Control
  29. Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation

    Botao Yu et al. · NeurIPS · 2022 · paper

    Introduces fine attention to selected structure-related bars and coarse summaries for other bars, allowing longer symbolic sequences and improved repetition structure.

    Cited in Why Long Songs Are Difficult
  30. Music Transformer

    Cheng-Zhi Anna Huang et al. · Google Brain and Magenta · 2018 · paper

    Original paper describing Transformer-based symbolic music generation with relative attention and long musical sequences.

    Cited in Prediction as the Basic Learning Task
  31. Music Transformer

    Cheng-Zhi Anna Huang et al. · Google Brain and Magenta · 2018 · paper

    Original paper describing autoregressive symbolic music generation with relative attention and minute-long piano event sequences.

    Cited in Autoregressive Music Models
  32. Music Transformer

    Cheng-Zhi Anna Huang et al. · Google Brain and Magenta · 2018 · paper

    Original paper describing relative attention for long symbolic sequences, minute-long piano generation, and motif-conditioned continuations.

    Cited in Why Long Songs Are Difficult
  33. Music Transformer: Generating Music with Long-Term Structure

    Google Magenta · Google Magenta · 2018 · website

    Official project page with generated piano examples and an accessible explanation of long-term musical structure.

    Cited in Prediction as the Basic Learning Task
  34. Music Transformer: Generating Music with Long-Term Structure

    Google Magenta · Google Magenta · 2018 · website

    Official project page with generated continuations and examples of long-term piano structure.

    Cited in Autoregressive Music Models
  35. Music Transformer: Generating Music with Long-Term Structure

    Google Magenta · Google Magenta · 2018 · website

    Official project page containing generated piano performances and an explanation of musical structure across several timescales.

    Cited in Why Long Songs Are Difficult
  36. MusicCaps Dataset Card

    Google · Google · 2023 · dataset

    Official dataset card for music clips paired with musician-written captions, supporting the connection between Chapter 2 caption quality and text-conditioned generation.

    Cited in Prediction as the Basic Learning Task
  37. MusicGen-Style

    Meta AI · Meta AudioCraft · 2024 · documentation

    Official documentation for combining a short style-reference excerpt with textual conditioning.

    Cited in Conditioning and Control
  38. MusicGen: Simple and Controllable Music Generation

    Meta AI · Meta AudioCraft · documentation

    Official implementation documentation describing autoregressive MusicGen token prediction and supported greedy, temperature, top-k, and top-p generation methods.

    Cited in Prediction as the Basic Learning Task
  39. MusicGen: Simple and Controllable Music Generation

    Meta AI · Meta AudioCraft · documentation

    Official implementation documentation describing conditional and unconditional generation, audio continuation, temperature, top-k, and top-p sampling.

    Cited in Autoregressive Music Models
  40. MusicGen: Simple and Controllable Music Generation

    Meta AI · Meta AudioCraft · documentation

    Official documentation for text generation, melody conditioning, and audio continuation in MusicGen.

    Cited in Conditioning and Control
  41. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original paper describing text-conditioned hierarchical sequence-to-sequence music generation using semantic and acoustic tokens.

    Cited in Hierarchical Generation
  42. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Original paper describing text-conditioned music generation and joint conditioning with hummed or whistled melodies.

    Cited in Conditioning and Control
  43. MusicLM: Generating Music From Text

    Andrea Agostinelli et al. · Google Research · 2023 · paper

    Describes hierarchical text-to-music generation at 24 kHz with reported consistency over several minutes.

    Cited in Why Long Songs Are Difficult
  44. Mustango: Toward Controllable Text-to-Music Generation

    Jan Melechovsky et al. · arXiv · 2023 · paper

    Introduces a diffusion-based text-to-music system with control over chords, beats, tempo, key, and general text descriptions.

    Cited in Conditioning and Control
  45. Neural Discrete Representation Learning

    Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · DeepMind · 2017 · paper

    Original VQ-VAE paper introducing learned discrete latent representations paired with autoregressive priors.

    Cited in Hierarchical Generation
  46. Noise2Music: Text-Conditioned Music Generation with Diffusion Models

    Qingqing Huang et al. · arXiv · 2023 · paper

    Introduces a cascade containing a text-conditioned diffusion generator and a diffusion cascader for high-fidelity 30-second music generation.

    Cited in Diffusion Models
  47. Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks

    Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · NeurIPS · 2015 · paper

    Original paper describing the mismatch between training on true previous tokens and inference from model-generated tokens, and how errors can accumulate along a sequence.

    Cited in Autoregressive Music Models
  48. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing a Transformer language model over several streams of compressed music tokens, conditioned on text or melody.

    Cited in Prediction as the Basic Learning Task
  49. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing one autoregressive Transformer over several streams of compressed music tokens using token interleaving patterns.

    Cited in Autoregressive Music Models
  50. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Original MusicGen paper describing text-conditioned and melody-conditioned generation over compressed music tokens.

    Cited in Conditioning and Control
  51. Simple and Controllable Music Generation

    Jade Copet et al. · Meta AI · 2023 · paper

    Describes autoregressive generation over several streams of compressed music tokens and provides context for the relationship between token rate and sequence length.

    Cited in Why Long Songs Are Difficult
  52. SoundStorm: Efficient Parallel Audio Generation

    Zalan Borsos et al. · Google Research · 2023 · paper

    Original paper replacing AudioLM's autoregressive acoustic stages with confidence-based parallel generation conditioned on semantic tokens.

    Cited in Hierarchical Generation
  53. SoundStream: An End-to-End Neural Audio Codec

    Neil Zeghidour et al. · Google Research · 2021 · paper

    Original paper describing a neural codec with residual vector quantisation used as the acoustic tokenizer in AudioLM and MusicLM.

    Cited in Hierarchical Generation
  54. Stable Audio Open

    Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · arXiv · 2024 · paper

    Describes an open-weights latent diffusion text-to-audio model and evaluates its audio autoencoder separately from the full generative system.

    Cited in Diffusion Models
  55. VampNet: Music Generation via Masked Acoustic Token Modeling

    Hugo Flores Garcia, Prem Seetharaman, Rithesh Kumar, and Bryan Pardo · ISMIR · 2023 · paper

    Introduces masked acoustic-token generation supporting synthesis, continuation, inpainting, outpainting, compression, and looping with variation.

    Cited in Conditioning and Control
  56. VampNet: Music Generation via Masked Acoustic Token Modeling

    Interactive Audio Lab · Interactive Audio Lab · 2023 · website

    Official project page with examples of VampNet continuation, inpainting, outpainting, and variation.

    Cited in Conditioning and Control
  57. WaveNet: A Generative Model for Raw Audio

    Aaron van den Oord et al. · Google DeepMind · 2016 · paper

    Original paper describing fully probabilistic autoregressive prediction of raw audio samples conditioned on previous samples, with speech and music examples.

    Cited in Prediction as the Basic Learning Task
  58. WaveNet: A Generative Model for Raw Audio

    Aaron van den Oord et al. · Google DeepMind · 2016 · paper

    Original paper describing probabilistic autoregressive waveform generation where each audio sample is conditioned on all previous samples.

    Cited in Autoregressive Music Models
  59. Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models

    Ziyu Wang, Lejun Min, and Gus Xia · arXiv · 2024 · paper

    Introduces hierarchical symbolic representations for whole-song form, phrases, cadences, chords, accompaniment, and notes, generated through a top-down cascade.

    Cited in Why Long Songs Are Difficult