Hands-On AI Science Series · In Production
Book cover: a glowing audio waveform rising through a wireframe network and equalizer bars into a luminous songbird in mid-flight, with the title Building Audio AI, From Waveforms to Generative Sound

Building Audio AI From Waveforms to Generative Sound

A practitioner's guide to digital signal processing, speech and audio understanding, generative sound, and real-time voice systems.

Alexander (Sasha) Apartsin, Ph.D. & Yehudit Aperstein, Ph.D.

Audio intelligence is the ability to convert pressure waves into meaning, action, and new sound. This book is one connected journey through the theories, models, and engineering practices for systems that hear, understand, speak, and deploy. It starts with the physics of sound and the signal-processing bedrock of sampling, filtering, and spectrograms, builds through classical audio ML and deep audio understanding, then moves into speech synthesis, music generation, and audio-language models, before closing with real-time voice agents, spatial and edge audio, and the evaluation, privacy, and deployment concerns that govern real systems.

10 parts 48 chapters 250+ sections 46 hands-on labs 7 appendices & a capstone

The Four-Verb Arc

Ten parts organized around four verbs; each stands on the one before it, from pressure waves to production systems.

How This Book Teaches

Five habits, kept in every chapter from the first sample to the last generated sound.

Worked Pipelines

Every chapter builds complete, runnable systems (a streaming keyword spotter, a CTC recognizer, a real-time voice assistant), never isolated snippets.

Library Shortcuts

After each from-scratch build, a shortcut callout shows the same task in a few lines of librosa, torchaudio, SpeechBrain, or Transformers, and names exactly what the library handles for you.

A Callout System

Pitfalls, math asides, practical industry examples, and cross-references are typeset as distinct boxes, so you can read deep or skim fast and never miss a trap.

Exercises & Labs

Each chapter closes with a hands-on lab that extends its worked pipelines, from quick checks to small projects you can put in a portfolio.

Classical Ideas Return Learned

Filterbanks become learnable frontends, masking becomes neural separation, vocoding becomes diffusion, and beamforming returns in neural spatial audio. One story, told twice.

The Hands-On AI Science Series

Building Audio AI is part of a family of connected books, each a deep, build-it-yourself guide to a major field of AI.

Hands-On AI Science is a series of in-depth guides to the major fields of artificial intelligence. Every book goes deep into the theory, models, and internals, covering the classical foundations and the most recent ideas, then shows you how to build each one in Python with the modern libraries and tools that get the job done. The writing stays plain and light (illustrations, analogies, mental models, worked examples, and a little fun) without trading away rigor or coverage. Each volume is self-contained and complete enough to anchor a full course on its subject.

Building Language AI

From Tokens to Agents.

Read online

Building Vision AI

From Pixels to Generative Models.

Read online

Building Audio AI

From Waveforms to Generative Sound.

You are here

Building Temporal AI

From Forecasting to Sequential Decision Making.

Read online

Building Scalable AI

From Big Data Algorithms to Distributed Intelligence.

Read online

Building Embodied AI

From Perception to Autonomous Action.

Read online

Building Agentic AI

From Goals to Autonomous Systems.

Read online

Building Discovery AI

From Vibe Coding to Autonomous Science.

Read online

Building Neuromorphic AI

From Spiking Neurons to Edge Intelligence.

Read online

Building Quantum AI

From Qubits to Quantum Machine Learning.

Read online

Building Tabular AI

From Structured Data to Decision Intelligence.

Read online