
A practitioner's guide to digital signal processing, speech and audio understanding, generative sound, and real-time voice systems.
Audio intelligence is the ability to convert pressure waves into meaning, action, and new sound. This book is one connected journey through the theories, models, and engineering practices for systems that hear, understand, speak, and deploy. It starts with the physics of sound and the signal-processing bedrock of sampling, filtering, and spectrograms, builds through classical audio ML and deep audio understanding, then moves into speech synthesis, music generation, and audio-language models, before closing with real-time voice agents, spatial and edge audio, and the evaluation, privacy, and deployment concerns that govern real systems.
Ten parts organized around four verbs; each stands on the one before it, from pressure waves to production systems.
The physical and computational nature of audio: sampling, quantization, filtering, the FFT and spectrograms, audio data engineering, and the classical features and baselines every system stands on.
Parts I–II · 9 chapters IINeural audio from CNNs to transformers and self-supervised encoders; speech recognition, speaker AI, and diarization; sound events, scenes, and monitoring across home, industry, health, and the wild.
Parts III–V · 16 chapters IIIThe synthesis stack: vocoders, text-to-speech, voice conversion, and streaming conversational speech; symbolic music, neural audio generation, text-to-audio models, and controllable sound design.
Parts VI–VII · 10 chapters IVAudio-language models and real-time voice agents; spatial audio, microphone arrays, and edge devices; evaluation, privacy, responsible audio AI, production serving, and a full capstone.
Parts VIII–X · 13 chaptersFive habits, kept in every chapter from the first sample to the last generated sound.
Every chapter builds complete, runnable systems (a streaming keyword spotter, a CTC recognizer, a real-time voice assistant), never isolated snippets.
After each from-scratch build, a shortcut callout shows the same task in a few lines of librosa, torchaudio, SpeechBrain, or Transformers, and names exactly what the library handles for you.
Pitfalls, math asides, practical industry examples, and cross-references are typeset as distinct boxes, so you can read deep or skim fast and never miss a trap.
Each chapter closes with a hands-on lab that extends its worked pipelines, from quick checks to small projects you can put in a portfolio.
Filterbanks become learnable frontends, masking becomes neural separation, vocoding becomes diffusion, and beamforming returns in neural spatial audio. One story, told twice.
Building Audio AI is part of a family of connected books, each a deep, build-it-yourself guide to a major field of AI.
Hands-On AI Science is a series of in-depth guides to the major fields of artificial intelligence. Every book goes deep into the theory, models, and internals, covering the classical foundations and the most recent ideas, then shows you how to build each one in Python with the modern libraries and tools that get the job done. The writing stays plain and light (illustrations, analogies, mental models, worked examples, and a little fun) without trading away rigor or coverage. Each volume is self-contained and complete enough to anchor a full course on its subject.
From Waveforms to Generative Sound.
You are here