Lunar Boom Learning
Sources
Course reference
Sources and citations
Each source remains connected to the lesson that uses it, even where a citation appears with different contextual notes.
Chapter 5: Evaluating, Releasing, and Governing the Model (121)
- 45 CFR 46: Protection of Human Subjects
Office for Human Research Protections · United States Department of Health and Human Services · regulation
Official United States regulations for covered human-subject research. Applicability depends on funding, institution, jurisdiction, and study characteristics.
Cited in Human Listening Tests → - ACE-Step: A Step Towards Music Generation Foundation Model
Junmin Gong, Sean Zhao, Sen Wang, Shengyuan Xu, and Jing Guo · arXiv · 2025 · paper
Primary research on a compressed diffusion and transformer music foundation model with multi-minute generation, temporal repainting, lyric editing, remixing, and specialized adaptation tasks.
Cited in Where AI Music Models Are Heading → - Adapting Frechet Audio Distance for Generative Music Evaluation
Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Emmanouilidou · IEEE ICASSP · 2024 · paper
Music-specific analysis of FAD sensitivity to sample count, embedding choice, reference-set quality, and perceptual correlation.
Cited in What Makes Generated Music Good? → - Adapting Fréchet Audio Distance for Generative Music Evaluation
Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Emmanouilidou · IEEE ICASSP · 2024 · paper
Music-specific study of FAD sample-size bias, audio-embedding choice, reference-set quality, and perceptual correlation.
Cited in Automated Evaluation → - AI Risk Management Framework Playbook: Manage
NIST · National Institute of Standards and Technology · government-guidance
Official guidance for deciding whether an AI system achieves its intended purpose and whether development or deployment should proceed.
Cited in What Makes Generated Music Good? → - AI Risk Management Framework Playbook: Manage
NIST · National Institute of Standards and Technology · government-guidance
Official guidance on post-deployment monitoring, external feedback, performance degradation, unexpected behavior, incidents, and risk response.
Cited in Turning a Model Into a Product → - AI Risk Management Framework Playbook: Map
NIST · National Institute of Standards and Technology · government-guidance
Official guidance for documenting intended purpose, users, deployment context, expectations, and impacts before selecting evaluation criteria.
Cited in What Makes Generated Music Good? → - AI Risk Management Framework Playbook: Measure
NIST · National Institute of Standards and Technology · government-guidance
Official guidance stating that evaluation approaches and metrics depend on purpose, audience, needs, and context, and should test whether a system is fit for purpose.
Cited in What Makes Generated Music Good? → - AI Risk Management Framework Playbook: Measure
NIST · National Institute of Standards and Technology · government-guidance
Official guidance emphasising that evaluation methods and metrics should be selected according to purpose, audience, context, and intended use.
Cited in Automated Evaluation → - AI Risk Management Framework Playbook: Measure
NIST · National Institute of Standards and Technology · government-guidance
Official guidance on selecting evaluation methods according to purpose, audience, context, intended use, and decision requirements.
Cited in Human Listening Tests → - Algorithms to Measure Audio Programme Loudness and True-Peak Audio Level
ITU Radiocommunication Sector · International Telecommunication Union · 2023 · standard
Official recommendation specifying algorithms for programme loudness and true-peak measurement.
Cited in Human Listening Tests → - Amazon SQS Visibility Timeout
AWS · Amazon Web Services · documentation
Official explanation of message invisibility during processing and redelivery when processing is not acknowledged before the timeout.
Cited in Turning a Model Into a Product → - An Early Look at the Possibilities as We Experiment with AI and Music
Lyor Cohen and Toni Reid · YouTube · 2023 · official-product-announcement
Official announcement of Dream Track, developed with participating artists who chose to collaborate and initially offered to selected creators.
Cited in Where AI Music Models Are Heading → - An Industrial-Strength Audio Search Algorithm
Avery Li-Chun Wang · International Society for Music Information Retrieval · 2003 · paper
Foundational description of a scalable spectral-landmark audio-fingerprinting system for identifying recordings from short, noisy, or compressed excerpts.
Cited in Similarity, Memorisation, and Attribution → - Arrange, Inpaint, and Refine: Steerable Long-term Music Audio Generation and Editing via Content-based Controls
Liwei Lin, Gus Xia, Yixiao Zhang, and Junyan Jiang · arXiv · 2024 · paper
Primary research adapting MusicGen for inpainting, score-conditioned arrangement, and track-conditioned refinement.
Cited in Where AI Music Models Are Heading → - Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
NIST · National Institute of Standards and Technology · 2024 · government-report
Official cross-sector guidance for managing generative-AI risks, including governance, pre-deployment testing, content provenance, monitoring, and incident disclosure.
Cited in Turning a Model Into a Product → - Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
NIST · National Institute of Standards and Technology · 2024 · government-guidance
Official cross-sector guidance for identifying, measuring, managing, and governing generative-AI risks across development and deployment.
Cited in Where AI Music Models Are Heading → - Asynchronous Inference
AWS · Amazon Web Services · documentation
Official example of queued asynchronous model serving for long processing times and large payloads, including scale-to-zero support.
Cited in Turning a Model Into a Product → - audiocraft.models.musicgen API Documentation
Meta AI · Meta AudioCraft · documentation
Official MusicGen generation API combining the required model components and exposing text and melody generation methods.
Cited in Turning a Model Into a Product → - AudioGen: Textually Guided Audio Generation
Felix Kreuk et al. · International Conference on Learning Representations · 2023 · paper
Original AudioGen paper using objective and subjective evaluation for text-conditioned audio generation.
Cited in Automated Evaluation → - AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Haohe Liu et al. · International Conference on Machine Learning · 2023 · paper
Original AudioLDM paper using CLAP representations and objective and subjective text-to-audio evaluation.
Cited in Automated Evaluation → - Automatic Evaluation of Speaker Similarity
Deja Kamil, Sanchez Ariadna, Roth Julian, and Cotescu Marius · arXiv · 2022 · preprint
Studies automatic speaker-similarity prediction using speaker embeddings and comparison with human perceptual ratings.
Cited in Similarity, Memorisation, and Attribution → - Autoscale an Asynchronous Endpoint
AWS · Amazon Web Services · documentation
Official guidance for autoscaling queued inference endpoints, including scaling to zero.
Cited in Turning a Model Into a Product → - Benchmarking Music Generation Models and Metrics via Human Preference Studies
Florian Grötschla, Ahmet Solak, Luca A. Lanzendörfer, and Roger Wattenhofer · arXiv · 2025 · paper
Large human-preference benchmark using generated songs from multiple systems, illustrating the difficulty of aligning automated music metrics with human judgement.
Cited in Where AI Music Models Are Heading → - C2PA and Content Credentials Explainer, Version 2.4
C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-guidance
Official explanation of Content Credentials, provenance, asset bindings, ingredients, AI disclosure, durable credentials, trust, privacy, and the limits of provenance assertions.
Cited in Similarity, Memorisation, and Attribution → - C2PA Technical Specification, Version 2.4
C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-standard
Current technical specification for cryptographically bound Content Credentials, including audio actions, ingredients, AI models, datasets, prompts, seeds, and generative-system asset types.
Cited in Similarity, Memorisation, and Attribution → - C2PA Technical Specification, Version 2.4
C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-standard
Current technical specification for cryptographically bound Content Credentials, asset actions, ingredients, audit logs, and AI and ML asset types.
Cited in Turning a Model Into a Product → - C2PA Technical Specification, Version 2.4
C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-standard
Current technical specification for Content Credentials, asset actions, ingredients, AI and ML asset types, and provenance assertions.
Cited in Where AI Music Models Are Heading → - Capability: Unit Economics
FinOps Foundation · FinOps Foundation · industry-guidance
Official FinOps guidance connecting cloud costs to product units, engineering decisions, pricing, and business outcomes.
Cited in Turning a Model Into a Product → - CLAP: Learning Audio Concepts From Natural Language Supervision
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · IEEE ICASSP · 2023 · paper
Original CLAP paper describing contrastive audio and language encoders trained to produce a joint embedding space.
Cited in What Makes Generated Music Good? → - CLAP: Learning Audio Concepts From Natural Language Supervision
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · IEEE ICASSP · 2023 · paper
Original CLAP paper describing contrastive audio and text encoders trained to produce a shared multimodal representation.
Cited in Automated Evaluation → - Cloud Cost Allocation
FinOps Foundation · FinOps Foundation · industry-guidance
Official FinOps guidance for assigning cloud costs to users, products, teams, projects, and other accountable groupings.
Cited in Turning a Model Into a Product → - Coarse Parallel Processing Using a Work Queue
Kubernetes contributors · Kubernetes · documentation
Official example in which parallel workers claim units of work from a task queue.
Cited in Turning a Model Into a Product → - CROWDMOS: An Approach for Crowdsourcing Mean Opinion Score Studies
Flavio Ribeiro, Dinei Florêncio, Cha Zhang, and Mike Seltzer · IEEE ICASSP · 2011 · paper
Original crowdsourced subjective-audio study proposing methods for identifying inaccurate responses and obtaining repeatable quality judgements outside a laboratory.
Cited in Human Listening Tests → - Datasheets for Datasets
Timnit Gebru et al. · Communications of the ACM · 2021 · paper
Original proposal for dataset documentation covering motivation, composition, collection, preprocessing, uses, distribution, maintenance, and limitations.
Cited in Similarity, Memorisation, and Attribution → - Datasheets for Datasets
Timnit Gebru et al. · Communications of the ACM · 2021 · paper
Original proposal for documenting dataset motivation, composition, collection, processing, uses, maintenance, and limitations.
Cited in Where AI Music Models Are Heading → - Download and Upload Objects with Presigned URLs
AWS · Amazon Web Services · documentation
Official guidance for time-limited upload and download access without exposing general storage credentials.
Cited in Turning a Model Into a Product → - ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck · Interspeech · 2020 · paper
Primary paper introducing a widely used speaker-embedding architecture for automatic speaker verification.
Cited in Similarity, Memorisation, and Attribution → - EnCodec Documentation
Meta AI · Meta AudioCraft · documentation
Official AudioCraft documentation for training and evaluating EnCodec.
Cited in Automated Evaluation → - Extracting Training Data from Large Language Models
Nicholas Carlini et al. · USENIX Security Symposium · 2021 · paper
Foundational demonstration that model querying can recover identifiable training examples. Included for the distinction between memorisation and extractability.
Cited in Similarity, Memorisation, and Attribution → - Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · Interspeech · 2019 · paper
Original Fréchet Audio Distance paper comparing distribution-level audio embeddings and human perception of audio distortions.
Cited in What Makes Generated Music Good? → - Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · Interspeech · 2019 · paper
Original Fréchet Audio Distance paper introducing distributional comparison of audio embeddings and validating it against artificial distortions and human perception.
Cited in Automated Evaluation → - General Methods for the Subjective Assessment of Sound Quality
ITU Radiocommunication Sector · International Telecommunication Union · 2019 · standard
Official recommendation covering general sound-quality assessment, grading scales, comparison tests, random presentation order, listening conditions, statistical analysis, and reporting.
Cited in Human Listening Tests → - Guidance for Artificial Intelligence and Machine Learning
C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-guidance
Official guidance for applying Content Credentials to datasets, models, software, and media used during training and inference.
Cited in Turning a Model Into a Product → - Guidance for Artificial Intelligence and Machine Learning, Version 2.4
C2PA · Coalition for Content Provenance and Authenticity · 2026 · technical-guidance
Official guidance for applying Content Credentials to AI datasets, software, models, fine-tuning, inference, outputs, versioning, and release.
Cited in Where AI Music Models Are Heading → - Guidance for the Selection of the Most Appropriate ITU-R Recommendation for Subjective Assessment of Sound Quality
ITU Radiocommunication Sector · International Telecommunication Union · standard
Official guidance for selecting a subjective audio-assessment method according to the test purpose and expected impairment.
Cited in Human Listening Tests → - High Fidelity Neural Audio Compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · Transactions on Machine Learning Research · 2023 · paper
Original EnCodec paper combining reconstruction objectives, perceptual losses, ablations, and MUSHRA listening tests for music and other audio domains.
Cited in Automated Evaluation → - How Robust is Audio Watermarking in Generative AI Models?
Yuxin Wen et al. · arXiv · 2025 · preprint
Evaluates recent audio-watermarking systems against a broad set of removal attacks and highlights the need to test claimed robustness under realistic transformations.
Cited in Similarity, Memorisation, and Attribution → - inference_mode
PyTorch contributors · PyTorch · documentation
Official documentation for disabling gradient-related tracking and selected autograd overhead during inference.
Cited in Turning a Model Into a Product → - Informed Consent FAQs
Office for Human Research Protections · United States Department of Health and Human Services · government-guidance
Official guidance on legally effective and documented informed consent for research covered by HHS requirements.
Cited in Human Listening Tests → - InvokeEndpointAsync
AWS · Amazon Web Services · api-documentation
Official API showing that asynchronous requests are enqueued and return result-location information rather than the inference result.
Cited in Turning a Model Into a Product → - ITU-R BS.1116: Methods for the Subjective Assessment of Small Impairments in Audio Systems
ITU Radiocommunication Sector · International Telecommunication Union · 2015 · standard
Official recommendation covering rigorous experimental design, listening conditions, subject selection, and reporting for critical audio assessment.
Cited in What Makes Generated Music Good? → - ITU-R BS.1283: Subjective Assessment of Sound Quality
ITU Radiocommunication Sector · International Telecommunication Union · 1997 · standard
Official guide explaining that subjective audio-assessment methods should be selected according to the purpose and scale of the expected impairments.
Cited in What Makes Generated Music Good? → - ITU-R BS.1534: Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems
ITU Radiocommunication Sector · International Telecommunication Union · 2015 · standard
Official recommendation defining the MUSHRA method and controlled subjective assessment principles for intermediate audio quality.
Cited in What Makes Generated Music Good? → - Jobs
Kubernetes contributors · Kubernetes · documentation
Official documentation for run-to-completion workloads, retries, parallelism, completion tracking, and failure policies.
Cited in Turning a Model Into a Product → - Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
Or Tal, Alon Ziv, Itai Gat, Felix Kreuk, and Yossi Adi · IEEE ICASSP · 2024 · paper
Primary JASCO paper combining text with temporally aligned chord, melody, drum, and full-mix controls.
Cited in Where AI Music Models Are Heading → - librosa Feature Extraction Documentation
librosa contributors · librosa · documentation
Official documentation for chroma, spectral, mel, tonal, onset, rhythm, and related audio features.
Cited in Automated Evaluation → - librosa.beat.beat_track
librosa contributors · librosa · documentation
Official documentation for dynamic-programming beat tracking using onset strength and tempo estimation.
Cited in Automated Evaluation → - Live Music Models
Lyria Team et al. · arXiv · 2025 · paper
Primary research introducing continuous real-time music streams with synchronized control and describing Magenta RealTime and Lyria RealTime.
Cited in Where AI Music Models Are Heading → - Long-form Music Generation with Latent Diffusion
Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · arXiv · 2024 · paper
Primary research reporting latent-diffusion music generation over long temporal contexts, with generated tracks of up to four minutes and forty-five seconds.
Cited in Where AI Music Models Are Heading → - Lyria 3
Google DeepMind · Google DeepMind · official-model-documentation
Official documentation describing Lyria 3, multimodal image conditioning, technical controls, and generation of cohesive tracks up to three minutes.
Cited in Where AI Music Models Are Heading → - Lyria RealTime
Google DeepMind · Google DeepMind · official-model-documentation
Official documentation describing real-time interactive music creation, control, and performance.
Cited in Where AI Music Models Are Heading → - Managing the Lifecycle of Objects
AWS · Amazon Web Services · documentation
Official documentation for transitioning objects to lower-cost classes and expiring objects through lifecycle rules.
Cited in Turning a Model Into a Product → - MelodySim: Measuring Melody-Aware Music Similarity for Plagiarism Detection
Tianze Lu et al. · arXiv · 2025 · preprint
Emerging research on melody-aware music similarity under transformations including note splitting, arpeggiation, track dropout, and re-instrumentation. Included as a developing method rather than an established universal standard.
Cited in Similarity, Memorisation, and Attribution → - MERT: Acoustic Music Understanding Model with Large-Scale Self-Supervised Training
Yizhi Li et al. · International Conference on Learning Representations · 2024 · paper
Original MERT paper describing a music-specialised self-supervised representation evaluated across fourteen music-understanding tasks.
Cited in Automated Evaluation → - Method for the Subjective Assessment of Intermediate Quality Level of Audio Systems
ITU Radiocommunication Sector · International Telecommunication Union · 2023 · standard
Official recommendation defining the MUSHRA method with multiple stimuli, hidden reference, anchors, controlled presentation, and subjective quality analysis.
Cited in Human Listening Tests → - Methods for the Subjective Assessment of Small Impairments in Audio Systems
ITU Radiocommunication Sector · International Telecommunication Union · 2015 · standard
Official recommendation for rigorous subjective evaluation of small audio impairments, including listening conditions, subject selection, training, test design, analysis, and reporting.
Cited in Human Listening Tests → - mir_eval Documentation
Colin Raffel et al. · mir_eval contributors · documentation
Official documentation for standardised evaluation functions used in music-information-retrieval research.
Cited in Automated Evaluation → - mir_eval: A Transparent Implementation of Common MIR Metrics
Colin Raffel et al. · International Society for Music Information Retrieval · 2014 · paper
Primary paper describing transparent implementations for beat, chord, melody, onset, pattern, segmentation, and other music-information-retrieval metrics.
Cited in Automated Evaluation → - mir_eval: A Transparent Implementation of Common MIR Metrics
Colin Raffel et al. · International Society for Music Information Retrieval · 2014 · paper
Primary source for transparent evaluation implementations covering melody, transcription, segmentation, rhythm, and other music-information-retrieval tasks.
Cited in Similarity, Memorisation, and Attribution → - mir_eval.melody Documentation
mir_eval contributors · mir_eval · documentation
Official documentation for evaluating voicing and pitch contours in melody extraction.
Cited in Similarity, Memorisation, and Attribution → - ML Model Registry
MLflow contributors · MLflow · documentation
Official documentation for registered model versions, aliases, tags, source runs, and deployment-oriented version management.
Cited in Turning a Model Into a Product → - MLflow Model Serving
MLflow contributors · MLflow · documentation
Official guidance on model signatures, environment capture, registry aliases, request validation, metadata, and serving.
Cited in Turning a Model Into a Product → - Model Cards for Model Reporting
Margaret Mitchell et al. · ACM Conference on Fairness, Accountability, and Transparency · 2019 · paper
Original proposal for model cards documenting intended uses, evaluation procedures, performance characteristics, limitations, and ethical considerations.
Cited in Similarity, Memorisation, and Attribution → - Model Cards for Model Reporting
Margaret Mitchell et al. · ACM Conference on Fairness, Accountability, and Transparency · 2019 · paper
Original proposal for structured model documentation covering intended use, evaluation, limitations, and relevant context.
Cited in Where AI Music Models Are Heading → - Model Hosting FAQs
AWS · Amazon Web Services · documentation
Official comparison of real-time, asynchronous, serverless, and batch inference patterns.
Cited in Turning a Model Into a Product → - Model Management
NVIDIA · NVIDIA · documentation
Official guidance on loading, unloading, adding, and removing model versions while serving traffic.
Cited in Turning a Model Into a Product → - Model Repository
NVIDIA · NVIDIA · documentation
Official model-repository layout and version-policy documentation.
Cited in Turning a Model Into a Product → - MuLan: A Joint Embedding of Music Audio and Natural Language
Qingqing Huang et al. · International Society for Music Information Retrieval · 2022 · paper
Original MuLan paper linking music audio with unconstrained natural-language descriptions for retrieval and zero-shot music understanding.
Cited in What Makes Generated Music Good? → - MuLan: A Joint Embedding of Music Audio and Natural Language
Qingqing Huang et al. · International Society for Music Information Retrieval · 2022 · paper
Original MuLan paper linking music audio with unconstrained natural-language descriptions for retrieval and zero-shot applications.
Cited in Automated Evaluation → - Music AI Sandbox, Now with New Features and Broader Access
Google DeepMind · Google DeepMind · 2025 · official-product-announcement
Official description of Lyria 2, Lyria RealTime, musician collaboration, continuous control, and SynthID watermarking.
Cited in Where AI Music Models Are Heading → - MusicGen Documentation
Meta AI · Meta AudioCraft · documentation
Official overview of MusicGen models, EnCodec tokenization, pretrained checkpoints, and inference use.
Cited in Turning a Model Into a Product → - MusicGen-Stem: Multi-stem Music Generation and Edition Through Autoregressive Modeling
Simon Rouard, Robin San Roman, Yossi Adi, and Axel Roebel · IEEE ICASSP · 2025 · paper
Primary research on parallel bass, drum, and other-stem generation, editing, and iterative source-conditioned composition.
Cited in Where AI Music Models Are Heading → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original MusicLM paper reporting separate evaluation of audio quality and text adherence and introducing MusicCaps.
Cited in What Makes Generated Music Good? → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · dataset-paper
Primary source for MusicCaps, described as approximately 5,500 music-text pairs with rich expert-written descriptions.
Cited in What Makes Generated Music Good? → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original MusicLM paper using automated and human measures of audio quality and text adherence and introducing the MusicCaps benchmark.
Cited in Automated Evaluation → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · dataset-paper
Primary source for MusicCaps, a benchmark of approximately 5,500 music-text pairs with detailed human-written descriptions.
Cited in Automated Evaluation → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original MusicLM paper describing a pairwise human test with two ten-second clips, a caption, graded A/B preference responses, and instructions focused specifically on text adherence.
Cited in Human Listening Tests → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original MusicLM paper containing a training-data memorisation analysis based on known training examples, prefixes, and continuation similarity.
Cited in Similarity, Memorisation, and Attribution → - MusicRL: Aligning Music Generation to Human Preferences
Geoffrey Cideron et al. · Proceedings of Machine Learning Research · 2024 · paper
Peer-reviewed ICML paper describing 300,000 pairwise music preferences and showing that text adherence and audio quality explain only part of human musical preference.
Cited in Human Listening Tests → - MusicRL: Aligning Music Generation to Human Preferences
Geoffrey Cideron et al. · Proceedings of Machine Learning Research · 2024 · paper
Peer-reviewed research using reward models and 300,000 pairwise preferences to align a MusicLM-based generator.
Cited in Where AI Music Models Are Heading → - Observability Primer
OpenTelemetry contributors · OpenTelemetry · documentation
Official introduction to observability and instrumentation for understanding distributed systems.
Cited in Turning a Model Into a Product → - OWASP Top 10 for Large Language Model Applications
OWASP contributors · OWASP · security-guidance
Community security guidance covering prompt injection, unsafe downstream output handling, excessive permissions, sensitive information, and related application risks.
Cited in Turning a Model Into a Product → - Proactive Detection of Voice Cloning with Localized Watermarking
Robin San Roman et al. · International Conference on Machine Learning · 2024 · paper
Introduces AudioSeal, a neural audio watermark with localised detection designed for identifying AI-generated speech regions.
Cited in Similarity, Memorisation, and Attribution → - Prompt Registry
MLflow contributors · MLflow · documentation
Official documentation for immutable prompt versions, aliases, comparison, reuse, A/B testing, and rollback.
Cited in Turning a Model Into a Product → - Quantifying Memorization Across Neural Language Models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang · International Conference on Learning Representations · 2023 · paper
Primary research connecting discoverable memorisation with model capacity, duplicated examples, and prompting context. MusicLM and MusicGen adapt related methodology to music.
Cited in Similarity, Memorisation, and Attribution → - Ragged Batching
NVIDIA · NVIDIA · documentation
Official guidance on batching requests with variable input shapes and the padding trade-off.
Cited in Turning a Model Into a Product → - Rank Analysis of Incomplete Block Designs: The Method of Paired Comparisons
Ralph Allan Bradley and Milton E. Terry · Biometrika · 1952 · paper
Foundational paper introducing the Bradley-Terry method for estimating rankings from pairwise comparison outcomes.
Cited in Human Listening Tests → - Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation
Baisen Wang, Chenxi Bao, and Qisong Han · arXiv · 2026 · preprint
Recent preprint exploring streaming latent generation, consistency distillation, and dynamic control for low-latency interaction. Included as emerging research rather than settled evidence.
Cited in Where AI Music Models Are Heading → - Secure Software Development Practices for Generative AI and Dual-Use Foundation Models
NIST · National Institute of Standards and Technology · 2024 · government-report
Official secure-development profile for producers and acquirers of AI models and AI systems.
Cited in Turning a Model Into a Product → - Signals
OpenTelemetry contributors · OpenTelemetry · documentation
Official definitions of traces, metrics, logs, baggage, and profiles.
Cited in Turning a Model Into a Product → - Simple and Controllable Music Generation
Jade Copet et al. · NeurIPS · 2023 · paper
Original MusicGen paper combining automated evaluation, human ratings of audio quality and text relevance, ablations, and memorisation testing.
Cited in What Makes Generated Music Good? → - Simple and Controllable Music Generation
Jade Copet et al. · NeurIPS · 2023 · paper
Original MusicGen paper combining text-alignment, audio-quality, human, diversity, ablation, and memorisation evaluations.
Cited in Automated Evaluation → - Simple and Controllable Music Generation
Jade Copet et al. · NeurIPS · 2023 · paper
Original MusicGen paper reporting separate human ratings for overall perceptual quality and text relevance, crowdsourced response filtering, several ratings per sample, confidence intervals, and loudness normalisation.
Cited in Human Listening Tests → - Simple and Controllable Music Generation
Jade Copet et al. · NeurIPS · 2023 · paper
Original MusicGen paper describing token-based memorisation analysis, training-example continuations, similarity thresholds, and human and automated evaluation.
Cited in Similarity, Memorisation, and Attribution → - Singer Identity Representation Learning using Self-Supervised Techniques
Research contributors · arXiv · 2024 · preprint
Investigates representations for singer identity and whether speech-trained speaker-recognition models transfer adequately to singing voice.
Cited in Similarity, Memorisation, and Attribution → - Sound Recordings Code Appendix on Artificial Intelligence
SAG-AFTRA and signatory record companies · SAG-AFTRA · 2024 · collective-bargaining-agreement
Contractual provisions for covered sound recordings, including definitions, consent requirements, compensation, and review of digital voice replicas. Scope is specific to the agreement.
Cited in Where AI Music Models Are Heading → - SoundLabs and Universal Music Group Announce Strategic Agreement to Offer Responsibly Trained AI Technology and Vocal Modeling Plug-In MicDrop to UMG Artists
Universal Music Group · Universal Music Group · 2024 · official-industry-announcement
Official announcement describing artist-specific vocal models trained on artists' own voice data, exclusive creative use, ownership control, and artistic approval.
Cited in Where AI Music Models Are Heading → - SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · IEEE/ACM Transactions on Audio, Speech, and Language Processing · 2021 · paper
Original SoundStream paper describing objective and subjective evaluation of an end-to-end neural codec across speech, music, and general audio.
Cited in Automated Evaluation → - SynthID
Google DeepMind · Google DeepMind · official-technical-documentation
Official description of watermarking and detection across AI-generated media, including audio generated through Lyria.
Cited in Where AI Music Models Are Heading → - The Belmont Report
National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research · United States Department of Health and Human Services · 1979 · government-guidance
Official ethical framework for research involving human participants, including respect for persons, beneficence, justice, and informed consent.
Cited in Human Listening Tests → - The Research Provisions
Information Commissioner's Office · Information Commissioner's Office · regulatory-guidance
Official data-protection guidance concerning personal-data processing for research, legal bases, safeguards, transparency, and participant rights.
Cited in Human Listening Tests → - Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez · International Society for Music Information Retrieval · 2024 · paper
Music-specific study comparing raw-audio similarity metrics and proposing the MiRA workflow for replication assessment.
Cited in What Makes Generated Music Good? → - Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez · International Society for Music Information Retrieval · 2024 · paper
Music-specific study evaluating raw-audio similarity metrics and proposing a model-independent replication-assessment workflow.
Cited in Automated Evaluation → - Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez · International Society for Music Information Retrieval · 2024 · paper
Introduces MiRA, a model-independent music-replication assessment workflow using several raw-audio similarity metrics and controlled replication experiments.
Cited in Similarity, Memorisation, and Attribution → - Traces
OpenTelemetry contributors · OpenTelemetry · documentation
Official explanation of end-to-end request traces and spans across services.
Cited in Turning a Model Into a Product → - Triton Inference Server Batchers
NVIDIA · NVIDIA · documentation
Official documentation for dynamic batching, queue delays, priorities, timeouts, and throughput-oriented request combination.
Cited in Turning a Model Into a Product → - Two Decades of Speaker Recognition Evaluation at the National Institute of Standards and Technology
Craig Greenberg et al. · National Institute of Standards and Technology · 2020 · government-paper
Official overview of systematic speaker-recognition and speaker-verification evaluation.
Cited in Similarity, Memorisation, and Attribution → - Universal Music Group and Udio Announce Udio's First Strategic Agreements for New Licensed AI Music Creation Platform
Universal Music Group · Universal Music Group · 2025 · official-industry-announcement
Official announcement describing a planned creation platform trained on authorized and licensed music, with fingerprinting, filtering, and protected platform controls.
Cited in Where AI Music Models Are Heading → - Using Dead-Letter Queues in Amazon SQS
AWS · Amazon Web Services · documentation
Official guidance for isolating messages that repeatedly fail processing.
Cited in Turning a Model Into a Product → - YuE: Scaling Open Foundation Models for Long-Form Music Generation
Ruibin Yuan et al. · arXiv · 2025 · paper
Primary research on long-form lyrics-to-song generation using track-decoupled prediction and structural progressive conditioning.
Cited in Where AI Music Models Are Heading →
Chapter 4: Training and Fine-Tuning a Music Model (113)
- A Gentle Introduction to torch.autograd
PyTorch contributors · PyTorch · documentation
Official explanation of forward propagation, backward propagation, automatic differentiation, computational graphs, and stored parameter gradients.
Cited in What Happens During Training? → - A Gentle Introduction to torch.autograd
PyTorch contributors · PyTorch · documentation
Official explanation of trainable parameters, gradient calculation, and optimisation foundations.
Cited in Training From Scratch Versus Fine-Tuning → - Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · International Conference on Learning Representations · 2015 · paper
Original paper introducing Adam and its adaptive first-order moment estimates.
Cited in What Happens During Training? → - Additional Guidance for the Training Pipeline
Google researchers and engineers · Google for Developers · documentation
Official guidance on profiling input pipelines, offline preprocessing, evaluation intervals, checkpoint selection, and experiment management.
Cited in Debugging Model Training → - Asynchronous Saving with Distributed Checkpoint
PyTorch contributors · PyTorch · documentation
Official recipe for reducing synchronous training interruption while writing distributed checkpoints.
Cited in Compute and Infrastructure → - Audio Conditioning for Music Generation via Discrete Bottleneck Features
Simon Rouard et al. · arXiv · 2024 · paper
Explores textual inversion using a pretrained text-to-music model and a jointly trained style-conditioning music model with quantised audio features.
Cited in Training From Scratch Versus Fine-Tuning → - Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning
Fang-Duo Tsai et al. · International Society for Music Information Retrieval · 2024 · paper
Introduces attention-based audio adapters for a pretrained AudioLDM2 model, with 22 million trainable parameters for timbre transfer, genre transfer, and accompaniment generation.
Cited in Training From Scratch Versus Fine-Tuning → - AudioCraft API Documentation
Meta AI · Meta AudioCraft · documentation
Official API documentation for AudioCraft training and inference components including MusicGen, AudioGen, and EnCodec.
Cited in Compute and Infrastructure → - AudioCraft Conditioning Documentation
Meta AI · Meta AudioCraft · documentation
Official documentation for text, waveform, joint-embedding, classifier-free guidance, and attribute-dropout conditioning.
Cited in Debugging Model Training → - AudioCraft Language Model API
Meta AI · Meta AudioCraft · documentation
Official API for MusicGen language modelling, sampling parameters, classifier-free guidance, condition dropout, and codebook generation.
Cited in Debugging Model Training → - AudioCraft MusicGen Model API
Meta AI · Meta AudioCraft · documentation
Official model interface for inference, text conditioning, melody conditioning, continuation, and generation parameters.
Cited in Training From Scratch Versus Fine-Tuning → - AudioCraft Training Documentation
Meta AI · Meta AudioCraft · documentation
Official documentation stating that AudioCraft provides training and inference code for MusicGen, AudioGen, and EnCodec.
Cited in Training From Scratch Versus Fine-Tuning → - AudioCraft Training Pipelines
Meta AI · Meta AudioCraft · documentation
Official documentation for AudioCraft's PyTorch training pipelines, Flashy, Dora experiment management, environment configuration, grids, clusters, and experiment workflows.
Cited in Compute and Infrastructure → - Autograd Mechanics
PyTorch contributors · PyTorch · documentation
Official technical documentation for reverse automatic differentiation and gradient computation.
Cited in What Happens During Training? → - Automatic Mixed Precision
PyTorch contributors · PyTorch · documentation
Official recipe for autocasting and gradient scaling in mixed-precision training.
Cited in Compute and Infrastructure → - Automatic Mixed Precision
PyTorch contributors · PyTorch · documentation
Official explanation of autocasting, gradient scaling, non-finite gradient handling, and mixed-precision optimiser steps.
Cited in Debugging Model Training → - Automatic Mixed Precision Examples
PyTorch contributors · PyTorch · documentation
Official examples covering autocasting, gradient scaling, gradient accumulation, clipping, and mixed-precision training.
Cited in What Happens During Training? → - Automatic Mixed Precision Examples
PyTorch contributors · PyTorch · documentation
Official examples for mixed precision, gradient accumulation, unscaling, clipping, and multiple models or losses.
Cited in Compute and Infrastructure → - Automatic Mixed Precision Examples
PyTorch contributors · PyTorch · documentation
Official examples for gradient accumulation, unscaling, clipping, and mixed-precision training.
Cited in Debugging Model Training → - Catastrophic Forgetting in the Context of Model Updates
Rich Harang and Hillary Sanders · arXiv · 2023 · paper
Compares retraining, fine-tuning, rehearsal, and regularisation approaches for updating models while retaining earlier performance.
Cited in Training From Scratch Versus Fine-Tuning → - Common Pitfalls and Recommended Practices
scikit-learn contributors · scikit-learn · documentation
Official guidance on data leakage, inconsistent preprocessing, test separation, and controlling randomness.
Cited in Overfitting and Memorisation → - Common Pitfalls and Recommended Practices
scikit-learn contributors · scikit-learn · documentation
Official guidance on data leakage, inconsistent preprocessing, split discipline, and random-state handling.
Cited in Debugging Model Training → - Cross-Validation: Evaluating Estimator Performance
scikit-learn contributors · scikit-learn · documentation
Official guidance distinguishing training, validation, and test data and warning that repeated hyperparameter decisions can leak information from the test set.
Cited in Designing a Small Experiment → - Cross-Validation: Evaluating Estimator Performance
scikit-learn contributors · scikit-learn · documentation
Official documentation distinguishing training, validation, and test use and explaining how cross-validation estimates generalisation.
Cited in Overfitting and Memorisation → - Cross-Validation: Evaluating Estimator Performance
scikit-learn contributors · scikit-learn · documentation
Official guidance for separating training, validation, and test use.
Cited in Debugging Model Training → - CrossEntropyLoss
PyTorch contributors · PyTorch · documentation
Official API documentation describing cross-entropy between input logits and targets.
Cited in What Happens During Training? → - CUDA Semantics
PyTorch contributors · PyTorch · documentation
Official documentation for CUDA behaviour, the caching allocator, allocated memory, reserved memory, and memory-management functions.
Cited in Compute and Infrastructure → - Data Loading Optimization in PyTorch
PyTorch contributors · PyTorch · documentation
Official tutorial covering data-loader workers, pinned memory, prefetching, persistent workers, and data-pipeline throughput.
Cited in Compute and Infrastructure → - Datasets and DataLoaders
PyTorch contributors · PyTorch · documentation
Official documentation for datasets, batching, shuffling, sampling, and iterable data loading.
Cited in What Happens During Training? → - Debugging Hangs with Flight Recorder
PyTorch contributors · PyTorch · documentation
Official tutorial for diagnosing distributed hangs with stack traces and collective-operation flight-recorder data.
Cited in Debugging Model Training → - Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · International Conference on Learning Representations · 2019 · paper
Original paper introducing decoupled weight decay for adaptive optimisers, commonly known as AdamW.
Cited in What Happens During Training? → - Deduplicating Training Data Makes Language Models Better
Katherine Lee et al. · Association for Computational Linguistics · 2022 · paper
Peer-reviewed language-model study on exact and near duplicate removal, memorised output, train-test overlap, and model efficiency. Used as cross-domain dataset evidence.
Cited in Overfitting and Memorisation → - Deduplicating Training Data Mitigates Privacy Risks in Language Models
Nikhil Kandpal, Eric Wallace, and Colin Raffel · Proceedings of Machine Learning Research · 2022 · paper
Peer-reviewed language-model study showing a strong relationship between duplicate count and regeneration frequency. Used here as cross-domain evidence, not as a music-specific result.
Cited in Overfitting and Memorisation → - Deep Learning Tuning Playbook
Google researchers and engineers · Google for Developers · documentation
Official practical guide to deep-learning experimentation, tuning, pipeline implementation, and systematic performance improvement.
Cited in Debugging Model Training → - Deep Learning Tuning Playbook FAQ
Google researchers and engineers · Google for Developers · documentation
Official guidance on identifying optimisation instability using learning-rate behaviour and gradient norms, with possible responses including warm-up and clipping.
Cited in Debugging Model Training → - DistributedDataParallel
PyTorch contributors · PyTorch · documentation
Official API documentation for model replication and gradient synchronisation across data-parallel processes.
Cited in Compute and Infrastructure → - Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · Journal of Machine Learning Research · 2014 · paper
Original peer-reviewed paper introducing dropout as a neural-network regularisation method.
Cited in Overfitting and Memorisation → - EnCodec Documentation
Meta AI · Meta AudioCraft · documentation
Official AudioCraft documentation for training and evaluating the neural audio codec used by MusicGen.
Cited in Compute and Infrastructure → - Exploring Adapter Design Tradeoffs for Low Resource Music Generation
Atharva Mehta, Shivam Chauhan, and Monojit Choudhury · arXiv · 2025 · paper
Music-generation study comparing adapter architectures and scales for adapting MusicGen and Mustango to underrepresented music traditions.
Cited in Training From Scratch Versus Fine-Tuning → - Extracting Training Data from Diffusion Models
Nicholas Carlini et al. · USENIX Security Symposium · 2023 · paper
Peer-reviewed image-diffusion study demonstrating training-data extraction through generation and filtering. Included as cross-domain evidence about extraction, not as direct evidence about music models.
Cited in Overfitting and Memorisation → - FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · NeurIPS · 2022 · paper
Introduces tiled exact attention that reduces memory reads and writes between GPU memory levels.
Cited in Compute and Infrastructure → - Generation or Replication: Auscultating Audio Latent Diffusion Models
Dimitrios Bralios et al. · IEEE ICASSP · 2024 · paper
Peer-reviewed audio study examining memorisation across training-set sizes, duplicated AudioCaps clips, and retrieval metrics for detecting replicated audio.
Cited in Overfitting and Memorisation → - Getting Started with DeviceMesh
PyTorch contributors · PyTorch · documentation
Official documentation for organising process groups and multi-dimensional distributed parallelism.
Cited in Compute and Infrastructure → - Getting Started with Distributed Checkpoint
PyTorch contributors · PyTorch · documentation
Official tutorial for saving and restoring state partitioned across distributed workers and changing trainer counts.
Cited in Compute and Infrastructure → - Getting Started with Distributed Data Parallel
PyTorch contributors · PyTorch · documentation
Official tutorial for multi-process and multi-machine data-parallel training.
Cited in Compute and Infrastructure → - Getting Started with Fully Sharded Data Parallel
PyTorch contributors · PyTorch · documentation
Official documentation describing the memory burden of model parameters, gradients, and optimiser states and how these can be sharded for large-scale training.
Cited in Training From Scratch Versus Fine-Tuning → - Getting Started with Fully Sharded Data Parallel
PyTorch contributors · PyTorch · documentation
Official tutorial explaining sharding of parameters, gradients, and optimiser states across workers.
Cited in Compute and Infrastructure → - GPU Performance Background User's Guide
NVIDIA · NVIDIA · documentation
Official explanation of GPU execution structure, memory hierarchy, and common performance limitations.
Cited in Compute and Infrastructure → - GroupKFold
scikit-learn contributors · scikit-learn · documentation
Official cross-validation method for domain-specific groups that should not overlap across folds.
Cited in Designing a Small Experiment → - GroupShuffleSplit
scikit-learn contributors · scikit-learn · documentation
Official grouped splitting tool that assigns unique groups rather than individual samples across partitions.
Cited in Designing a Small Experiment → - How to Save Memory by Fusing the Optimizer Step into the Backward Pass
PyTorch contributors · PyTorch · documentation
Official tutorial illustrating memory used by parameters, activations, gradients, optimiser states, and optimiser intermediates during training.
Cited in Compute and Infrastructure → - Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
Or Tal et al. · International Society for Music Information Retrieval · 2024 · paper
Introduces JASCO for global text conditioning together with local symbolic and audio controls using a new flow-matching architecture and conditioning method.
Cited in Training From Scratch Versus Fine-Tuning → - Large Scale Transformer Model Training with Tensor Parallel
PyTorch contributors · PyTorch · documentation
Official tutorial for combining tensor parallelism with fully sharded training for large Transformer models.
Cited in Compute and Infrastructure → - LoRA Conceptual Guide
Hugging Face contributors · Hugging Face · documentation
Official conceptual guidance for LoRA matrices, target modules, and parameter-efficient adaptation.
Cited in Training From Scratch Versus Fine-Tuning → - LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu et al. · International Conference on Learning Representations · 2022 · paper
Original LoRA paper describing frozen pretrained weights with trainable low-rank matrix updates.
Cited in Training From Scratch Versus Fine-Tuning → - MAGNeT Training and Fine-Tuning Documentation
Meta AI · Meta AudioCraft · documentation
Official documentation showing checkpoint-based fine-tuning and warning that configurations and changed layers must remain compatible.
Cited in Training From Scratch Versus Fine-Tuning → - Mixed Precision Training
Paulius Micikevicius et al. · International Conference on Learning Representations · 2018 · paper
Foundational paper describing lower-precision training, higher-precision weight updates, and loss scaling.
Cited in Compute and Infrastructure → - Mosaic: Memory Profiling for PyTorch
PyTorch contributors · PyTorch · documentation
Official tutorial for analysing peak GPU memory, allocation categories, activation checkpointing savings, and memory imbalance.
Cited in Compute and Infrastructure → - Music Transformer
Cheng-Zhi Anna Huang et al. · Google Brain and Magenta · 2018 · paper
Original paper describing relative attention for long symbolic music sequences, expressive piano generation, motif continuation, and comparisons with recurrent approaches.
Cited in Designing a Small Experiment → - Music Transformer: Generating Music with Long-Term Structure
Google Magenta · Google Magenta · 2018 · website
Official project page explaining the event representation and the contrast between recurrent hidden-state compression and Transformer access to earlier events.
Cited in Designing a Small Experiment → - MusicGen Documentation
Meta AI · Meta AudioCraft · documentation
Official overview of the MusicGen model, codec-token representation, and AudioCraft implementation.
Cited in What Happens During Training? → - MusicGen Documentation
Meta AI · Meta AudioCraft · documentation
Official overview of MusicGen architecture, pretrained checkpoints, generation modes, and codec-token configuration.
Cited in Training From Scratch Versus Fine-Tuning → - MusicGen Documentation
Meta AI · Meta AudioCraft · documentation
Official documentation describing MusicGen's 32 kHz EnCodec tokenizer, four codebooks at 50 Hz, generation modes, and model configuration.
Cited in Compute and Infrastructure → - MusicGen Documentation
Meta AI · Meta AudioCraft · documentation
Official documentation for MusicGen architecture, EnCodec token streams, conditioning, and generation.
Cited in Debugging Model Training → - MusicGenSolver API Documentation
Meta AI · Meta AudioCraft · documentation
Official implementation showing MusicGen batch preparation, tokenisation, masking, per-codebook cross-entropy, backpropagation, gradient scaling, clipping, optimiser updates, scheduling, and metric logging.
Cited in What Happens During Training? → - MusicGenSolver API Documentation
Meta AI · Meta AudioCraft · documentation
Official implementation documentation for MusicGen training, validation, generation, model summaries, metrics, and optimisation.
Cited in Compute and Infrastructure → - MusicGenSolver API Documentation
Meta AI · Meta AudioCraft · documentation
Official implementation showing metadata and audio checks, condition preparation, codec encoding, padding masks, per-codebook loss, backward propagation, gradient scaling, clipping, optimiser updates, finite-loss checks, generation, and audio decoding.
Cited in Debugging Model Training → - NVIDIA Deep Learning Performance Guide
NVIDIA · NVIDIA · documentation
Official guidance on GPU execution, batch sizing, memory-limited operations, and deep-learning performance.
Cited in Compute and Infrastructure → - Optimizing Model Parameters
PyTorch contributors · PyTorch · documentation
Official tutorial covering loss functions, optimisers, zeroing gradients, backward propagation, optimiser steps, and training and validation loops.
Cited in What Happens During Training? → - Overcoming Catastrophic Forgetting in Neural Networks
James Kirkpatrick et al. · Proceedings of the National Academy of Sciences · 2017 · paper
Foundational study of catastrophic forgetting during sequential neural-network training and a regularisation method for protecting important earlier parameters.
Cited in Training From Scratch Versus Fine-Tuning → - Parameter-Efficient Fine-Tuning
Hugging Face contributors · Hugging Face · documentation
Official PEFT documentation covering efficient adaptation without updating all base-model parameters.
Cited in Training From Scratch Versus Fine-Tuning → - Parameter-Efficient Transfer Learning for Music Foundation Models
Yiwei Ding and Alexander Lerch · arXiv · 2024 · paper
Music-specific comparison of adapter-based, prompt-based, and reparameterisation-based transfer methods against probing, fine-tuning, and small models trained from scratch.
Cited in Training From Scratch Versus Fine-Tuning → - Parameter-Efficient Transfer Learning for NLP
Neil Houlsby et al. · International Conference on Machine Learning · 2019 · paper
Original paper introducing compact task-specific adapter modules while keeping the original network fixed.
Cited in Training From Scratch Versus Fine-Tuning → - Performance RNN: Generating Music with Expressive Timing and Dynamics
Ian Simon and Sageev Oore · Google Magenta · 2017 · website
Official Magenta explanation of an LSTM-based recurrent model using event-based MIDI to generate polyphonic piano performances with expressive timing and dynamics.
Cited in Designing a Small Experiment → - Performance Tuning Guide
PyTorch contributors · PyTorch · documentation
Official guidance for improving performance across data loading, memory, precision, and model execution.
Cited in Compute and Infrastructure → - Performance Tuning Guide
PyTorch contributors · PyTorch · documentation
Official performance guidance noting that anomaly detection, profilers, and gradcheck are debugging tools rather than normal training settings.
Cited in Debugging Model Training → - pretty_midi Documentation
Colin Raffel and contributors · Colin Raffel · documentation
Official documentation for parsing, inspecting, modifying, synthesising, and extracting information from MIDI files.
Cited in Designing a Small Experiment → - Reproducibility
PyTorch contributors · PyTorch · documentation
Official guidance on random seeds, deterministic behaviour, DataLoader worker seeding, and the limits of reproducibility across releases and platforms.
Cited in Designing a Small Experiment → - Reproducibility
PyTorch contributors · PyTorch · documentation
Official guidance on seeds, deterministic algorithms, data-loader randomness, and limits to exact reproducibility across releases and platforms.
Cited in Compute and Infrastructure → - Reproducibility
PyTorch contributors · PyTorch · documentation
Official guidance on random seeds, deterministic behaviour, data-loader randomness, and reproducibility limitations.
Cited in Debugging Model Training → - Saving and Loading Models
PyTorch contributors · PyTorch · documentation
Official guidance for saving model state and general training checkpoints.
Cited in What Happens During Training? → - Saving and Loading Models
Matthew Inkawhich and PyTorch contributors · PyTorch · documentation
Official guidance for saving weights, general checkpoints, optimiser state, and restoring models.
Cited in Compute and Infrastructure → - Saving and Loading Models
PyTorch contributors · PyTorch · documentation
Official guidance for preserving model, optimiser, scheduler, and training state.
Cited in Debugging Model Training → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing a Transformer language model trained over several streams of compressed music tokens.
Cited in What Happens During Training? → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing a music language model trained over compressed EnCodec token streams.
Cited in Training From Scratch Versus Fine-Tuning → - Simple and Controllable Music Generation
Jade Copet et al. · NeurIPS · 2023 · paper
Original MusicGen paper. Its memorisation experiment tests exact and 80 percent partial matches in the first codec stream for five-second continuations from 20,000 sampled training examples.
Cited in Overfitting and Memorisation → - Simple and Controllable Music Generation
Jade Copet et al. · NeurIPS · 2023 · paper
Original MusicGen paper describing a single-stage Transformer over multiple streams of compressed music tokens.
Cited in Compute and Infrastructure → - Simple and Controllable Music Generation
Jade Copet et al. · NeurIPS · 2023 · paper
Original MusicGen paper describing conditional autoregressive modelling over several streams of compressed music tokens.
Cited in Debugging Model Training → - The Lakh MIDI Dataset v0.1
Colin Raffel · Colin Raffel · dataset
Official project page describing 176,581 unique MIDI files, including 45,129 matched and aligned with Million Song Dataset entries.
Cited in Designing a Small Experiment → - The MAESTRO Dataset
Curtis Hawthorne et al. · Google Magenta · dataset
Official dataset page describing roughly 200 hours of aligned piano audio and MIDI, expressive performance information, metadata, predefined splits, and licence terms.
Cited in Designing a Small Experiment → - torch.autograd
PyTorch contributors · PyTorch · documentation
Official automatic-differentiation documentation relevant to gradients and backward-pass state.
Cited in Compute and Infrastructure → - torch.autograd
PyTorch contributors · PyTorch · documentation
Official automatic-differentiation documentation, including anomaly detection and gradient-checking tools.
Cited in Debugging Model Training → - torch.isfinite
PyTorch contributors · PyTorch · documentation
Official API for identifying finite and non-finite tensor elements.
Cited in Debugging Model Training → - torch.nn.Dropout
PyTorch contributors · PyTorch · documentation
Official API documentation explaining that dropout randomly zeros activations during training and behaves differently during evaluation.
Cited in Overfitting and Memorisation → - torch.nn.Module
PyTorch contributors · PyTorch · documentation
Official module API containing forward and backward hook mechanisms for inspecting model computations.
Cited in Debugging Model Training → - torch.nn.utils.clip_grad_norm_
PyTorch contributors · PyTorch · documentation
Official documentation for calculating and clipping the total norm of parameter gradients.
Cited in What Happens During Training? → - torch.nn.utils.clip_grad_norm_
PyTorch contributors · PyTorch · documentation
Official API for calculating and clipping the total gradient norm.
Cited in Debugging Model Training → - torch.optim
PyTorch contributors · PyTorch · documentation
Official documentation for optimisation algorithms and learning-rate schedulers.
Cited in What Happens During Training? → - torch.profiler
PyTorch contributors · PyTorch · documentation
Official profiler documentation for tracing CPU and CUDA operations, shapes, stacks, and memory.
Cited in Compute and Infrastructure → - torch.profiler
PyTorch contributors · PyTorch · documentation
Official profiler for collecting performance and memory information during training and inference.
Cited in Debugging Model Training → - torch.Tensor.register_hook
PyTorch contributors · PyTorch · documentation
Official API for registering callbacks executed when tensor gradients are calculated.
Cited in Debugging Model Training → - torch.utils.checkpoint
PyTorch contributors · PyTorch · documentation
Official documentation explaining activation checkpointing as a compute-for-memory trade-off.
Cited in Compute and Infrastructure → - torch.utils.data
PyTorch contributors · PyTorch · documentation
Official dataset, sampler, batching, worker, prefetch, and persistent-worker documentation.
Cited in Compute and Infrastructure → - torch.utils.data
PyTorch contributors · PyTorch · documentation
Official documentation for datasets, batching, sampling, workers, prefetching, and data loading.
Cited in Debugging Model Training → - Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
Roser Batlle-Roca, Wei-Hsiang Liao, Xavier Serra, Yuki Mitsufuji, and Emilia Gómez · International Society for Music Information Retrieval · 2024 · paper
Peer-reviewed music-specific work proposing the MiRA assessment approach and evaluating several raw-audio similarity metrics in controlled replication tests.
Cited in Overfitting and Memorisation → - Transfer Learning for Computer Vision Tutorial
Sasank Chilamkurthy and PyTorch contributors · PyTorch · documentation
Official PyTorch tutorial contrasting full fine-tuning with freezing a pretrained network and training only a final layer.
Cited in Training From Scratch Versus Fine-Tuning → - Underfitting Versus Overfitting
scikit-learn contributors · scikit-learn · documentation
Official example illustrating underfitting, suitable complexity, overfitting, and cross-validated error.
Cited in Overfitting and Memorisation → - Understanding CUDA Memory Usage
PyTorch contributors · PyTorch · documentation
Official guidance for recording and visualising CUDA memory snapshots and investigating allocation patterns.
Cited in Compute and Infrastructure → - Validation Curves: Plotting Scores to Evaluate Models
scikit-learn contributors · scikit-learn · documentation
Official documentation explaining training and validation curves, underfitting, overfitting, and how additional data can affect generalisation.
Cited in Overfitting and Memorisation → - Visualizing Models, Data, and Training with TensorBoard
PyTorch contributors · PyTorch · documentation
Official tutorial for recording and visualising training metrics, images, model graphs, and embeddings.
Cited in Compute and Infrastructure → - What Is torch.nn Really?
PyTorch contributors · PyTorch · documentation
Official tutorial distinguishing training and evaluation modes and calculating validation loss.
Cited in What Happens During Training? → - ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · SC20 · 2020 · paper
Foundational work on eliminating replicated parameter, gradient, and optimiser-state memory in distributed training.
Cited in Compute and Infrastructure → - Zeroing Out Gradients in PyTorch
PyTorch contributors · PyTorch · documentation
Official explanation of gradient accumulation and the need to clear gradients between ordinary training steps.
Cited in What Happens During Training? →
Chapter 2: Building the Musical Training Dataset (51)
- About CC Licenses
Creative Commons · Creative Commons · documentation
Official overview of Creative Commons licence elements and the permissions and conditions attached to standard CC licences.
Cited in Licensing, Consent, and Provenance → - Announcing AudioSet: A Dataset for Audio Event Research
Google Research · Google Research · 2017 · website
Official announcement describing AudioSet as more than two million ten-second excerpts labelled with hundreds of sound-event categories.
Cited in What Is a Training Example? → - Audio Based Disambiguation of Music Genre Tags
Romain Hennequin, Jimena Royo-Letelier, and Manuel Moussallam · arXiv · 2018 · paper
Research showing how audio-based genre embeddings can identify duplicate labels, translate between tag systems, and recover genre relationships.
Cited in Metadata and Captioning → - Audio Resampling
Caroline Chen and Moto Hira · PyTorch · documentation
Official tutorial on bandlimited audio resampling, filter parameters, roll-off, aliasing, and efficient reuse of resampling kernels.
Cited in Collecting and Cleaning Audio → - Audio Set: An Ontology and Human-Labeled Dataset for Audio Events
Jort F. Gemmeke et al. · Google Research · 2017 · paper
Original paper describing a large-scale dataset of manually labelled audio events organised through an ontology.
Cited in What Is a Training Example? → - Audio Set: An Ontology and Human-Labeled Dataset for Audio Events
Jort F. Gemmeke et al. · Google Research · 2017 · paper
Original paper describing a manually annotated, multi-label audio-event dataset organised through a structured ontology.
Cited in Metadata and Captioning → - Audio Set: An Ontology and Human-Labeled Dataset for Audio Events
Jort F. Gemmeke et al. · Google Research · 2017 · paper
Original paper introducing AudioSet, its sound-event ontology, and its large multi-label collection with highly uneven class frequencies.
Cited in Dataset Balance and Bias → - Bias beyond Borders: Global Inequalities in AI-Generated Music
Ahmet Solak, Florian Grötschla, Luca A. Lanzendörfer, and Roger Wattenhofer · arXiv · 2025 · paper
Introduces the globally balanced GlobalDISCO evaluation dataset and reports disparities across regions, languages, and mainstream versus geographically niche genres.
Cited in Dataset Balance and Bias → - Can MusicGen Create Training Data for MIR Tasks?
Nadine Kroher, Helena Cuesta, and Aggelos Pikrakis · arXiv · 2023 · paper
Early experiment using MusicGen to create genre-conditioned synthetic training data for a music information retrieval classifier.
Cited in Dataset Balance and Bias → - Chromaprint
Lukas Lalinsky and contributors · AcoustID · documentation
Official repository describing Chromaprint as a compact audio fingerprint system designed for near-identical full-file identification and duplicate audio detection.
Cited in Collecting and Cleaning Audio → - Copyright and Artificial Intelligence, Part 3: Generative AI Training
U.S. Copyright Office · U.S. Copyright Office · 2025 · documentation
Official report describing the legal and policy debate around copyrighted works in generative AI training and the fact-specific, jurisdiction-sensitive nature of the analysis.
Cited in Licensing, Consent, and Provenance → - Copyright Registration of Musical Compositions and Sound Recordings
U.S. Copyright Office · U.S. Copyright Office · 2020 · documentation
Official circular explaining that musical compositions and sound recordings are separate works and identifying the different material and authors associated with each.
Cited in Licensing, Consent, and Provenance → - Creative Commons Frequently Asked Questions
Creative Commons · Creative Commons · 2025 · documentation
Official FAQ explaining that CC licences do not automatically clear third-party publicity, privacy, personality, trademark, patent, or other rights outside their scope.
Cited in Licensing, Consent, and Provenance → - Croissant Format Specification 1.1
MLCommons Croissant Working Group · MLCommons · 2026 · documentation
Machine-learning dataset metadata standard supporting licence fields and PROV-O relationships at dataset, resource, record, field, and value level.
Cited in Licensing, Consent, and Provenance → - DALI: A Large Dataset of Synchronized Audio, Lyrics and Notes
Gabriel Meseguer-Brocal, Alice Cohen-Hadria, and Geoffroy Peeters · arXiv · 2019 · paper
Original paper introducing 5,358 audio tracks with time-aligned lyrics and vocal melody notes at four levels of granularity.
Cited in What Is a Training Example? → - Data Leakage in Cross-Modal Retrieval Training: A Case Study
Benno Weck and Xavier Serra · Music Technology Group, Universitat Pompeu Fabra · 2023 · paper
Audio-specific study finding duplicates across SoundDesc train and evaluation data, showing overly optimistic retrieval performance, and proposing corrected splits.
Cited in Collecting and Cleaning Audio → - Dataset Balancing Can Hurt Model Performance
R. Channing Moore, Daniel P. W. Ellis, Eduardo Fonseca, Shawn Hershey, Aren Jansen, and Manoj Plakal · arXiv · 2023 · paper
Audio-specific study showing that balancing AudioSet improved one public evaluation score while hurting a separate evaluation set and did not clearly improve rare-class performance.
Cited in Dataset Balance and Bias → - Datasheets for Datasets
Timnit Gebru et al. · Communications of the ACM · 2021 · paper
Foundational framework proposing documentation of dataset motivation, composition, collection, processing, uses, distribution, and maintenance.
Cited in Dataset Balance and Bias → - Datasheets for Datasets
Timnit Gebru et al. · Communications of the ACM · 2021 · paper
Foundational dataset-documentation framework covering motivation, composition, collection, processing, distribution, maintenance, uses, and limitations.
Cited in Licensing, Consent, and Provenance → - Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset
Curtis Hawthorne et al. · Google Magenta · 2018 · paper
Original paper introducing paired piano audio and MIDI with fine alignment for transcription, symbolic composition, and waveform synthesis.
Cited in What Is a Training Example? → - Ethical Dimensions of Music Information Retrieval Technology
Andre Holzapfel et al. · Transactions of the International Society for Music Information Retrieval · 2018 · paper
Music-specific discussion of cultural, social, data, and algorithmic biases in music information retrieval research and technology.
Cited in Dataset Balance and Bias → - ffprobe Documentation
FFmpeg Project · FFmpeg · documentation
Official documentation for inspecting media containers and streams, selecting audio streams, checking usable stream configuration, and producing structured output.
Cited in Collecting and Cleaning Audio → - Jukebox: A Generative Model for Music
Prafulla Dhariwal et al. · OpenAI · 2020 · paper
Original Jukebox paper describing music generation conditioned on artist, genre, and unaligned lyrics.
Cited in What Is a Training Example? → - Leveraging Knowledge Bases and Parallel Annotations for Music Genre Translation
Elena V. Epure, Anis Khlif, and Romain Hennequin · arXiv · 2019 · paper
Research addressing diversity and subjectivity across genre tag systems and methods for translating annotations between taxonomies.
Cited in Metadata and Captioning → - librosa.effects.split
librosa development team · librosa · documentation
Official documentation for identifying non-silent intervals in mono or multichannel audio.
Cited in Collecting and Cleaning Audio → - librosa.effects.trim
librosa development team · librosa · documentation
Official documentation for trimming leading and trailing silence using a threshold below a reference level.
Cited in Collecting and Cleaning Audio → - librosa.resample
librosa development team · librosa · documentation
Official documentation for converting a time series from an original sample rate to a target sample rate.
Cited in Collecting and Cleaning Audio → - LP-MusicCaps
SeungHeon Doh and contributors · KAIST Music and Audio Computing Lab · 2023 · documentation
Official project repository documenting tag-to-caption generation and audio-to-caption training.
Cited in Metadata and Captioning → - LP-MusicCaps: LLM-Based Pseudo Music Captioning
SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam · International Society for Music Information Retrieval · 2023 · paper
Original paper generating large-scale captions from music tags and evaluating caption quality through correct attributes and incorrect attributes.
Cited in Metadata and Captioning → - Missing Melodies: AI Music Generation and its Nearly Complete Omission of the Global South
Atharva Mehta, Shivam Chauhan, and Monojit Choudhury · arXiv · 2024 · paper
Large review of more than one million dataset hours and over 200 papers examining geographic and cultural representation in music-generation research.
Cited in Dataset Balance and Bias → - Music Classification Annotations
Music Technology Group · Music Technology Group, Universitat Pompeu Fabra · dataset
Official annotation documentation describing raw labels from three annotators and a clean subset containing agreed labels while excluding unmatched and instrumental answers where relevant.
Cited in Metadata and Captioning → - Music for All: Exploring Multicultural Representations in Music Generation Models
Atharva Mehta, Shivam Chauhan, Amirbek Djanibekov, Atharva Kulkarni, Gus Xia, and Monojit Choudhury · arXiv · 2025 · paper
Music-specific study measuring genre and cultural underrepresentation in generation datasets and testing parameter-efficient adaptation for Hindustani classical and Turkish makam music.
Cited in Dataset Balance and Bias → - MusicCaps Dataset Card
Google · Google · 2023 · dataset
Official dataset card describing 5,521 records and fields including source video ID, segment timing, AudioSet labels, aspect lists, and musician-written captions.
Cited in What Is a Training Example? → - MusicCaps Dataset Card
Google · Google · 2023 · dataset
Official dataset card describing 5,521 ten-second music examples with AudioSet labels, aspect lists, and free-text captions written by musicians.
Cited in Metadata and Captioning → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original MusicLM paper introducing text-conditioned music generation and the MusicCaps dataset of about 5.5 thousand music-text pairs.
Cited in What Is a Training Example? → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original MusicLM paper describing MusicCaps, expert-written captions, aspect lists, text conditioning, and caption adherence evaluation.
Cited in Metadata and Captioning → - Position: Measure Dataset Diversity, Don't Just Claim It
Sabina Tomkins et al. · arXiv · 2024 · paper
Argues that claims about dataset diversity need explicit definitions, dimensions, and measurements rather than broad unsupported descriptions.
Cited in Dataset Balance and Bias → - PROV-O: The PROV Ontology
W3C Provenance Working Group · World Wide Web Consortium · 2013 · documentation
W3C Recommendation providing classes and relationships for representing and exchanging provenance information across systems.
Cited in Licensing, Consent, and Provenance → - Recommendation ITU-R BS.1770-5: Algorithms to Measure Audio Programme Loudness and True-Peak Audio Level
ITU Radiocommunication Sector · International Telecommunication Union · 2023 · documentation
Official recommendation defining objective programme loudness and true-peak measurement algorithms.
Cited in Collecting and Cleaning Audio → - Regulation (EU) 2016/679, General Data Protection Regulation
European Parliament and Council of the European Union · European Union · 2016 · documentation
Official EU regulation defining consent and, where processing relies on consent, requiring that it be demonstrable and withdrawable under the stated conditions.
Cited in Licensing, Consent, and Provenance → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing a music generation model conditioned on text descriptions or melodic features.
Cited in What Is a Training Example? → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing full-length music standardised to 32 kHz, mono downmixing in the main setup, and random 30-second training crops.
Cited in Collecting and Cleaning Audio → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing music generation conditioned on textual descriptions or melodic features.
Cited in Metadata and Captioning → - Stable Audio Open
Zach Evans et al. · Stability AI · 2024 · paper
Original paper describing an open text-to-audio model producing stereo audio at 44.1 kHz.
Cited in Collecting and Cleaning Audio → - Stable Audio Tools Dataset Documentation
Stability AI · Stability AI · documentation
Official dataset documentation describing loading and resampling, optional LUFS normalisation, silence handling, padding and cropping, padding masks, and channel conversion.
Cited in Collecting and Cleaning Audio → - The MAESTRO Dataset
Google Magenta · Google Magenta · 2018 · dataset
Official dataset page describing paired audio and MIDI, key velocities, pedal positions, descriptive metadata, and alignment of about 3 milliseconds.
Cited in What Is a Training Example? → - The MAESTRO Dataset
Google Magenta · Google Magenta · 2018 · dataset
Official dataset page stating that MAESTRO is available under the Creative Commons Attribution Non-Commercial Share-Alike 4.0 licence.
Cited in Licensing, Consent, and Provenance → - The MTG-Jamendo Dataset
Dmitry Bogdanov et al. · Music Technology Group, Universitat Pompeu Fabra · 2019 · dataset
Official dataset page describing more than 55,000 tracks and 195 uploader-provided genre, instrument, and mood or theme tags.
Cited in Metadata and Captioning → - The MTG-Jamendo Dataset
Dmitry Bogdanov et al. · Music Technology Group, Universitat Pompeu Fabra · 2019 · dataset
Official dataset page describing more than 55,000 tracks, 195 genre, instrument, and mood or theme tags, uploader-provided labels, published splits, and tag distributions.
Cited in Dataset Balance and Bias → - Using Audio Fingerprinting for Duplicate Detection and Thumbnail Generation
Christopher J. C. Burges, John C. Platt, and Soumya Jana · Microsoft Research · 2005 · paper
Research describing audio fingerprinting for detecting duplicate clips even when files differ in compression quality or duration.
Cited in Collecting and Cleaning Audio → - Using CC-Licensed Works for AI Training
Creative Commons · Creative Commons · 2025 · documentation
Official guidance explaining that copyright treatment of AI training varies and discussing conservative handling of attribution, share-alike, noncommercial, and no-derivatives licence elements.
Cited in Licensing, Consent, and Provenance →
Chapter 1: Turning Music Into Data (31)
- About MIDI Part 3: MIDI Messages
The MIDI Association · The MIDI Association · documentation
Official overview of note on, note off, velocity, program change, pitch bend, aftertouch, and control change messages.
Cited in Three Ways to Represent Music → - About MIDI Part 4: MIDI Files
The MIDI Association · The MIDI Association · documentation
Official explanation that Standard MIDI Files store timed instructions and events rather than recorded sound, with playback produced by a connected sound engine.
Cited in Three Ways to Represent Music → - Audio Set: An Ontology and Human-Labeled Dataset for Audio Events
Jort F. Gemmeke et al. · Google Research · 2017 · paper
Original paper introducing the AudioSet ontology and large-scale human-labelled dataset used for audio event recognition.
Cited in Audio Embeddings → - AudioLM: A Language Modeling Approach to Audio Generation
Zalan Borsos et al. · Google Research · 2022 · paper
Original AudioLM paper describing audio generation as language modelling over discrete audio tokens and combining semantic and acoustic token types.
Cited in Audio Tokens and Neural Codecs → - CLAP: Learning Audio Concepts From Natural Language Supervision
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang · Microsoft Research · 2022 · paper
Original CLAP paper describing separate audio and text encoders trained contrastively in a shared multimodal embedding space.
Cited in Audio Embeddings → - Core Audio Essentials
Apple Developer Documentation · 2017 · documentation
Official overview of audio data formats, including sample rate, bits per channel, and channels per frame.
Cited in What Is Sound? → - Digitizing Audio
Adobe · 2021 · documentation
Official documentation covering microphones, digital sampling, sample rate, bit depth, amplitude values, channels, and audio file size.
Cited in What Is Sound? → - Displaying Audio in the Waveform Editor
Adobe · 2021 · documentation
Official documentation explaining waveform and spectral displays, their axes, amplitude peaks, frequency components, colour intensity, and the time-frequency resolution tradeoff.
Cited in Three Ways to Represent Music → - Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset
Curtis Hawthorne et al. · Google Magenta · 2019 · paper
Original MAESTRO paper presenting aligned audio and note events for transcription, symbolic composition, and waveform synthesis.
Cited in Three Ways to Represent Music → - EnCodec: High Fidelity Neural Audio Compression
Meta Research · Meta Research · documentation
Official repository documenting available 24 kHz and 48 kHz EnCodec models and supported bitrate settings.
Cited in Audio Tokens and Neural Codecs → - GANSynth: Adversarial Neural Audio Synthesis
Jesse Engel et al. · Google Magenta · 2019 · paper
Original paper describing neural audio synthesis in the spectral domain using log magnitudes and instantaneous frequencies.
Cited in Three Ways to Represent Music → - High Fidelity Neural Audio Compression
Alexandre Defossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · Meta AI · 2022 · paper
Original EnCodec paper describing a real-time encoder-decoder codec with a quantized latent space and evaluations across speech, music, mono, and stereo audio.
Cited in Audio Tokens and Neural Codecs → - High-Fidelity Audio Compression with Improved RVQGAN
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · Descript · 2023 · paper
Original Descript Audio Codec paper describing high-fidelity universal audio compression into discrete codes using improved residual vector quantization and reconstruction training.
Cited in Audio Tokens and Neural Codecs → - How Do We Hear?
National Institute on Deafness and Other Communication Disorders · 2022 · website
Official explanation of how sound waves produce vibrations that are converted into electrical signals in the auditory system.
Cited in What Is Sound? → - MuLan: A Joint Embedding of Music Audio and Natural Language
Qingqing Huang et al. · Google Research · 2022 · paper
Original MuLan paper describing a two-tower joint music and text embedding model for retrieval, tagging, transfer learning, and zero-shot tasks.
Cited in Audio Embeddings → - Music Transformer
Cheng-Zhi Anna Huang et al. · Google Brain · 2018 · paper
Original paper describing a Transformer model for long symbolic piano performances, continuations, and melody-conditioned accompaniment.
Cited in Three Ways to Represent Music → - Music Transformer: Generating Music with Long-Term Structure
Google Magenta · Google Magenta · 2018 · website
Official project overview and generated examples for Music Transformer.
Cited in Three Ways to Represent Music → - MusicGen: Simple and Controllable Music Generation
Meta AI · Meta AudioCraft · documentation
Official implementation documentation describing MusicGen's 32 kHz EnCodec tokenizer and model configuration.
Cited in What Is Sound? → - MusicGen: Simple and Controllable Music Generation
Meta AI · Meta AudioCraft · documentation
Official implementation documentation describing MusicGen's 32 kHz EnCodec tokenizer, four codebooks, and 50 Hz token rate.
Cited in Audio Tokens and Neural Codecs → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original MusicLM paper explaining how MuLan music embeddings were used during training and MuLan text embeddings were used as conditioning during generation.
Cited in Audio Embeddings → - scipy.signal.spectrogram
SciPy · documentation
Official SciPy documentation describing a spectrogram as a view of changing frequency content over time and explaining its segmented Fourier transform parameters.
Cited in Three Ways to Represent Music → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing compressed discrete music representations, common music sample rates, and mono and stereo generation.
Cited in What Is Sound? → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing one Transformer language model operating over several streams of compressed discrete music tokens.
Cited in Audio Tokens and Neural Codecs → - Sound Classification with YAMNet
TensorFlow · 2024 · documentation
Official TensorFlow Hub tutorial describing YAMNet, its AudioSet classes, audio input requirements, and its score, embedding, and spectrogram outputs.
Cited in Audio Embeddings → - SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour et al. · Google Research · 2021 · paper
Original paper describing an end-to-end encoder, residual vector quantizer, and decoder for speech, music, and general audio, including variable bitrate operation.
Cited in Audio Tokens and Neural Codecs → - The MAESTRO Dataset
Google Magenta · Google Magenta · 2018 · dataset
Official dataset page describing about 200 hours of paired piano audio and MIDI aligned to about 3 milliseconds, including velocity and pedal data.
Cited in Three Ways to Represent Music → - The Science of Sound
NASA · website
Official overview of sound as a longitudinal wave and the relationship of amplitude and frequency to loudness and pitch.
Cited in What Is Sound? → - Transfer Learning with YAMNet for Environmental Sound Classification
TensorFlow · 2024 · documentation
Official TensorFlow tutorial describing YAMNet's 1,024-dimensional embeddings and their use as pretrained features for a smaller audio classifier.
Cited in Audio Embeddings → - UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Leland McInnes, John Healy, and James Melville · arXiv · 2018 · paper
Original UMAP paper describing a scalable dimensionality reduction method used to project high-dimensional data for visualisation.
Cited in Audio Embeddings → - WAVEFORMATEX Structure
Microsoft Learn · 2021 · documentation
Official format documentation defining channel count, samples per second, and bits per sample for waveform audio.
Cited in What Is Sound? → - WaveNet: A Generative Model for Raw Audio
Aaron van den Oord et al. · Google DeepMind · 2016 · paper
Original paper introducing an autoregressive model that generates raw audio waveforms one sample at a time and includes experiments on music.
Cited in Three Ways to Represent Music →
Chapter 3: How an AI Model Generates Music (59)
- Audio Conditioning for Music Generation via Discrete Bottleneck Features
Simon Rouard, Yossi Adi, Jade Copet, Axel Roebel, and Alexandre Défossez · arXiv · 2024 · paper
Introduces audio-based style conditioning, joint text and audio control, and separate guidance over the two condition types.
Cited in Conditioning and Control → - AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Haohe Liu et al. · ICML · 2023 · paper
Introduces text-to-audio generation through latent diffusion, an audio autoencoder, and aligned CLAP audio and text embeddings.
Cited in Diffusion Models → - AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Haohe Liu et al. · ICML · 2023 · paper
Describes text-guided latent audio diffusion using aligned audio and text embeddings.
Cited in Conditioning and Control → - AudioLM Examples
Google Research · Google Research · 2022 · website
Official project page with speech and piano continuation examples and the AudioLM abstract.
Cited in Prediction as the Basic Learning Task → - AudioLM Examples
Google Research · Google Research · 2022 · website
Official project page containing speech and piano audio-continuation examples.
Cited in Hierarchical Generation → - AudioLM: A Language Modeling Approach to Audio Generation
Zalan Borsos et al. · Google Research · 2022 · paper
Original paper mapping audio to discrete tokens and treating audio generation as language modelling, including piano continuations from audio prompts.
Cited in Prediction as the Basic Learning Task → - AudioLM: A Language Modeling Approach to Audio Generation
Zalan Borsos et al. · Google Research · 2022 · paper
Original paper describing hierarchical autoregressive language models over semantic and neural codec tokens for long-term audio consistency and acoustic quality.
Cited in Autoregressive Music Models → - AudioLM: A Language Modeling Approach to Audio Generation
Zalan Borsos et al. · Google Research · 2022 · paper
Original paper describing three-stage hierarchical generation with semantic, coarse acoustic, and fine acoustic tokens.
Cited in Hierarchical Generation → - Classifier-Free Diffusion Guidance
Jonathan Ho and Tim Salimans · arXiv · 2022 · paper
Introduces guidance based on combined conditional and unconditional diffusion predictions without requiring a separate classifier.
Cited in Diffusion Models → - Classifier-Free Diffusion Guidance
Jonathan Ho and Tim Salimans · arXiv · 2022 · paper
Introduces guidance based on combined conditional and unconditional predictions.
Cited in Conditioning and Control → - CrossEntropyLoss
PyTorch contributors · PyTorch · documentation
Official documentation for computing cross-entropy loss between model logits and target classes.
Cited in Prediction as the Basic Learning Task → - Denoising Diffusion Implicit Models
Jiaming Song, Chenlin Meng, and Stefano Ermon · ICLR · 2021 · paper
Introduces non-Markovian diffusion sampling processes that use the same training objective while supporting faster generation with fewer sampling steps.
Cited in Diffusion Models → - Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · NeurIPS · 2020 · paper
Foundational paper describing a Gaussian forward corruption process, learned reverse denoising process, and noise-prediction training objective.
Cited in Diffusion Models → - DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · ICLR · 2021 · paper
Original paper describing conditional and unconditional waveform generation by converting white noise into structured audio through diffusion.
Cited in Diffusion Models → - Fast Timing-Conditioned Latent Audio Diffusion
Zach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley, and Jordi Pons · arXiv · 2024 · paper
Describes Stable Audio's text and timing conditioning for long-form, variable-length stereo generation.
Cited in Conditioning and Control → - High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · CVPR · 2022 · paper
Foundational latent diffusion paper explaining diffusion in the compressed representation of a pretrained autoencoder to reduce computational requirements.
Cited in Diffusion Models → - Improved Denoising Diffusion Probabilistic Models
Alexander Quinn Nichol and Prafulla Dhariwal · ICML · 2021 · paper
Introduces improvements including learned reverse variances and an alternative noise schedule, with substantially fewer sampling passes in tested settings.
Cited in Diffusion Models → - JASCO: Joint Audio and Symbolic Conditioning
Meta AI · Meta AudioCraft · 2024 · documentation
Official AudioCraft documentation for temporally controlled generation using text, symbolic, and audio conditions.
Cited in Conditioning and Control → - Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
Or Tal, Alon Ziv, Itai Gat, Felix Kreuk, and Yossi Adi · arXiv · 2024 · paper
Introduces JASCO, combining global text with time-aligned chord, melody, drum, and full-mix controls.
Cited in Conditioning and Control → - Jukebox: A Generative Model for Music
Prafulla Dhariwal et al. · OpenAI · 2020 · paper
Original paper describing a multi-scale VQ-VAE and autoregressive Transformer priors for generating music with singing over multiple minutes.
Cited in Autoregressive Music Models → - Jukebox: A Generative Model for Music
Prafulla Dhariwal et al. · OpenAI · 2020 · paper
Original paper describing multi-scale VQ-VAE compression, a top-level autoregressive prior, and middle and bottom upsampling models.
Cited in Hierarchical Generation → - Jukebox: A Generative Model for Music
Prafulla Dhariwal et al. · OpenAI · 2020 · paper
Describes artist, genre, and unaligned lyric conditioning for raw-audio song generation.
Cited in Conditioning and Control → - Jukebox: A Generative Model for Music
Prafulla Dhariwal et al. · OpenAI · 2020 · paper
Describes a multi-scale VQ-VAE and autoregressive Transformer priors for raw-audio music generation with reported coherence over multiple minutes.
Cited in Why Long Songs Are Difficult → - Long-form Music Generation with Latent Diffusion
Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · arXiv · 2024 · paper
Describes a timing-conditioned diffusion Transformer operating on a highly downsampled continuous audio latent for long-form text-to-music generation.
Cited in Diffusion Models → - Long-form Music Generation with Latent Diffusion
Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · arXiv · 2024 · paper
Describes a diffusion Transformer operating at a latent rate of 21.5 positions per second, trained on long temporal contexts and generating music up to 4 minutes and 45 seconds.
Cited in Why Long Songs Are Difficult → - MELONS: Generating Melody with Long-Term Structure Using Transformers and Structure Graph
Yi Zou et al. · ICASSP · 2022 · paper
Introduces a graph representation of bar-level melody relationships and separates structure generation from structure-conditioned melody generation.
Cited in Why Long Songs Are Difficult → - Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion
Flavio Schneider, Ojasv Kamal, Zhijing Jin, and Bernhard Schölkopf · arXiv · 2023 · paper
Describes a two-stage system using a diffusion autoencoder and text-conditioned latent diffusion for long stereo music generation.
Cited in Diffusion Models → - MuseCoco: Generating Symbolic Music from Text
Peiling Lu et al. · arXiv · 2023 · paper
Introduces text-to-symbolic-music generation using explicit musical attributes as an intermediate control representation.
Cited in Conditioning and Control → - Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation
Botao Yu et al. · NeurIPS · 2022 · paper
Introduces fine attention to selected structure-related bars and coarse summaries for other bars, allowing longer symbolic sequences and improved repetition structure.
Cited in Why Long Songs Are Difficult → - Music Transformer
Cheng-Zhi Anna Huang et al. · Google Brain and Magenta · 2018 · paper
Original paper describing Transformer-based symbolic music generation with relative attention and long musical sequences.
Cited in Prediction as the Basic Learning Task → - Music Transformer
Cheng-Zhi Anna Huang et al. · Google Brain and Magenta · 2018 · paper
Original paper describing autoregressive symbolic music generation with relative attention and minute-long piano event sequences.
Cited in Autoregressive Music Models → - Music Transformer
Cheng-Zhi Anna Huang et al. · Google Brain and Magenta · 2018 · paper
Original paper describing relative attention for long symbolic sequences, minute-long piano generation, and motif-conditioned continuations.
Cited in Why Long Songs Are Difficult → - Music Transformer: Generating Music with Long-Term Structure
Google Magenta · Google Magenta · 2018 · website
Official project page with generated piano examples and an accessible explanation of long-term musical structure.
Cited in Prediction as the Basic Learning Task → - Music Transformer: Generating Music with Long-Term Structure
Google Magenta · Google Magenta · 2018 · website
Official project page with generated continuations and examples of long-term piano structure.
Cited in Autoregressive Music Models → - Music Transformer: Generating Music with Long-Term Structure
Google Magenta · Google Magenta · 2018 · website
Official project page containing generated piano performances and an explanation of musical structure across several timescales.
Cited in Why Long Songs Are Difficult → - MusicCaps Dataset Card
Google · Google · 2023 · dataset
Official dataset card for music clips paired with musician-written captions, supporting the connection between Chapter 2 caption quality and text-conditioned generation.
Cited in Prediction as the Basic Learning Task → - MusicGen-Style
Meta AI · Meta AudioCraft · 2024 · documentation
Official documentation for combining a short style-reference excerpt with textual conditioning.
Cited in Conditioning and Control → - MusicGen: Simple and Controllable Music Generation
Meta AI · Meta AudioCraft · documentation
Official implementation documentation describing autoregressive MusicGen token prediction and supported greedy, temperature, top-k, and top-p generation methods.
Cited in Prediction as the Basic Learning Task → - MusicGen: Simple and Controllable Music Generation
Meta AI · Meta AudioCraft · documentation
Official implementation documentation describing conditional and unconditional generation, audio continuation, temperature, top-k, and top-p sampling.
Cited in Autoregressive Music Models → - MusicGen: Simple and Controllable Music Generation
Meta AI · Meta AudioCraft · documentation
Official documentation for text generation, melody conditioning, and audio continuation in MusicGen.
Cited in Conditioning and Control → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original paper describing text-conditioned hierarchical sequence-to-sequence music generation using semantic and acoustic tokens.
Cited in Hierarchical Generation → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Original paper describing text-conditioned music generation and joint conditioning with hummed or whistled melodies.
Cited in Conditioning and Control → - MusicLM: Generating Music From Text
Andrea Agostinelli et al. · Google Research · 2023 · paper
Describes hierarchical text-to-music generation at 24 kHz with reported consistency over several minutes.
Cited in Why Long Songs Are Difficult → - Mustango: Toward Controllable Text-to-Music Generation
Jan Melechovsky et al. · arXiv · 2023 · paper
Introduces a diffusion-based text-to-music system with control over chords, beats, tempo, key, and general text descriptions.
Cited in Conditioning and Control → - Neural Discrete Representation Learning
Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · DeepMind · 2017 · paper
Original VQ-VAE paper introducing learned discrete latent representations paired with autoregressive priors.
Cited in Hierarchical Generation → - Noise2Music: Text-Conditioned Music Generation with Diffusion Models
Qingqing Huang et al. · arXiv · 2023 · paper
Introduces a cascade containing a text-conditioned diffusion generator and a diffusion cascader for high-fidelity 30-second music generation.
Cited in Diffusion Models → - Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · NeurIPS · 2015 · paper
Original paper describing the mismatch between training on true previous tokens and inference from model-generated tokens, and how errors can accumulate along a sequence.
Cited in Autoregressive Music Models → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing a Transformer language model over several streams of compressed music tokens, conditioned on text or melody.
Cited in Prediction as the Basic Learning Task → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing one autoregressive Transformer over several streams of compressed music tokens using token interleaving patterns.
Cited in Autoregressive Music Models → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Original MusicGen paper describing text-conditioned and melody-conditioned generation over compressed music tokens.
Cited in Conditioning and Control → - Simple and Controllable Music Generation
Jade Copet et al. · Meta AI · 2023 · paper
Describes autoregressive generation over several streams of compressed music tokens and provides context for the relationship between token rate and sequence length.
Cited in Why Long Songs Are Difficult → - SoundStorm: Efficient Parallel Audio Generation
Zalan Borsos et al. · Google Research · 2023 · paper
Original paper replacing AudioLM's autoregressive acoustic stages with confidence-based parallel generation conditioned on semantic tokens.
Cited in Hierarchical Generation → - SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour et al. · Google Research · 2021 · paper
Original paper describing a neural codec with residual vector quantisation used as the acoustic tokenizer in AudioLM and MusicLM.
Cited in Hierarchical Generation → - Stable Audio Open
Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons · arXiv · 2024 · paper
Describes an open-weights latent diffusion text-to-audio model and evaluates its audio autoencoder separately from the full generative system.
Cited in Diffusion Models → - VampNet: Music Generation via Masked Acoustic Token Modeling
Hugo Flores Garcia, Prem Seetharaman, Rithesh Kumar, and Bryan Pardo · ISMIR · 2023 · paper
Introduces masked acoustic-token generation supporting synthesis, continuation, inpainting, outpainting, compression, and looping with variation.
Cited in Conditioning and Control → - VampNet: Music Generation via Masked Acoustic Token Modeling
Interactive Audio Lab · Interactive Audio Lab · 2023 · website
Official project page with examples of VampNet continuation, inpainting, outpainting, and variation.
Cited in Conditioning and Control → - WaveNet: A Generative Model for Raw Audio
Aaron van den Oord et al. · Google DeepMind · 2016 · paper
Original paper describing fully probabilistic autoregressive prediction of raw audio samples conditioned on previous samples, with speech and music examples.
Cited in Prediction as the Basic Learning Task → - WaveNet: A Generative Model for Raw Audio
Aaron van den Oord et al. · Google DeepMind · 2016 · paper
Original paper describing probabilistic autoregressive waveform generation where each audio sample is conditioned on all previous samples.
Cited in Autoregressive Music Models → - Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
Ziyu Wang, Lejun Min, and Gus Xia · arXiv · 2024 · paper
Introduces hierarchical symbolic representations for whole-song form, phrases, cadences, chords, accompaniment, and notes, generated through a top-down cascade.
Cited in Why Long Songs Are Difficult →