Ad blocker App image

25 Studio Presets, Zero Servers: On-Device Audio Processing with AVAudioEngine

25 Studio Presets, Zero Servers: On-Device Audio Processing with AVAudioEngine

By Ameya Chorghade • Sr. iOS Developer at Pixster Studio • 8 October 2026 • 6 min read

Zero Servers, 25 Presets

For our upcoming video app (not launched yet), we built 25 one-tap audio presets, from Denoise to Cathedral, that run entirely on the iPhone. No server, no third-party DSP library, no upload.


This post walks through how a single AVAudioEngine chain built from Apple’s own audio units powers every one of them. The Audio Studio screen runs every chip in its grid through the same processing chain; only the numbers change.

The Problem

Most creators record video on their phone, in a kitchen, a car or a cafe. The audio comes out boomy, hissy or uneven, and that is often the first thing viewers notice.


We wanted a one-tap fix: pick a preset, hear the difference immediately, export. We also set three constraints:

  • No upload. Users’ audio never leaves the device. That is better for privacy and avoids server cost per minute of audio.

  • Instant preview. Switching presets must feel like flipping a switch, not waiting for a render.

  • No heavy dependencies. No third-party DSP SDK to license, update or debug.

The Approach: One Chain, Many Numbers

The key design decision was to build exactly one processing chain and make every preset nothing more than a set of numbers fed into it.


Apple ships everything we needed inside AVFoundation and AudioToolbox:

  • AVAudioUnitEQ for the high-pass filter and parametric EQ

  • AVAudioUnitEffect wrapping kAudioUnitSubType_DynamicsProcessor for compression and noise gating

  • AVAudioUnitReverb with its factory room presets

  • AVAudioEngine’s main mixer for final gain


Because the chain never changes shape, adding preset number 26 is a data change, not a code change. It also means preview and export can share the same logic.

How the Chain Works

Audio flows through five stages in a fixed order, then out to either the speaker (preview) or a file (export).

  1. High-pass filter removes everything below a cutoff. Around 30–60 Hz removes only rumble, 100 Hz removes muddiness while keeping the body of a voice, and 400–500 Hz removes all bass for a small-speaker sound.

  2. Parametric EQ boosts or cuts around chosen frequencies. Roughly: 50–150 Hz is body, 200–400 Hz is mud and boxiness, 1–3 kHz is voice clarity, 5–8 kHz is crispness, and 10 kHz and up is air and hiss.

  3. Dynamics processor does two jobs. The compressor squashes anything above its threshold so loud and soft words sit closer together. The expander pushes anything below its threshold down, which works as a noise gate between words.

  4. Reverb adds a room, from a subtle small room to a full cathedral.

  5. Gain normalizes the loudest peak to −1 dBFS and adds the preset’s output boost.


Order matters. Filtering before compression means the compressor reacts to the voice, not to rumble. Adding reverb after compression keeps the room sound from being squashed.

Presets Are Just Data

Every preset is a plain struct. Here is Denoise, our default preset, trimmed for readability:

struct EQBandConfig {

    let frequency: Float // Hz

    let gain: Float // dB

    let bandwidth: Float // octaves

}


static let denoise = AudioPreset(

    id: “denoise”, name: “Denoise”,

    eqBands: [

        EQBandConfig(frequency: 200, gain: -3.5, bandwidth: 0.35), // sub-vocal rumble

        EQBandConfig(frequency: 320, gain: -2.5, bandwidth: 0.25), // narrow notch on room boxiness

        EQBandConfig(frequency: 2200, gain: 3.5, bandwidth: 0.9), // vocal presence

        EQBandConfig(frequency: 7500, gain: -5.0, bandwidth: 0.7), // gentle roll-off above consonants

        EQBandConfig(frequency: 9500, gain: -12.0, bandwidth: 1.0) // hiss territory

    ],

    compressionThreshold: -16, compressionHeadroom: 5,

    expansionRatio: 10.0, expansionThreshold: -32, // hard gate between words

    attackTime: 0.001, releaseTime: 0.22,

    reverbPreset: nil, reverbMix: 0,

    normalizeVolume: true, highPassFreq: 100, outputGain: 3.0

)

The same chain produces very different results depending on those numbers:

Preset

High-pass

EQ shape

Compression

Gate

Reverb

Denoise

100 Hz

Cuts lows and hiss, lifts 2.2 kHz

Moderate

10:1, hard

None

Podcast

130 Hz

Big +7 dB presence boost at 2.5 kHz

Firm

3.5:1

None

Telephone

400 Hz

Boost around 1.2 kHz, −18 dB above 6 kHz

Heavy

4:1

None

Cathedral

60 Hz

Gentle warmth and air

Light

2:1

Cathedral, 65% wet

Denoise and Telephone use identical code. The only difference is where the numbers put the energy.

Plotting the combined high-pass and EQ response makes the difference visible. Denoise stays close to flat through the voice range, then cuts about 14 dB around 8 to 9 kHz, where hiss lives. Podcast pushes presence around 2.6 kHz. Telephone carves a narrow band around 1.2 kHz, and Bass Boost lifts everything below 130 Hz.

Two Engines, One Chain

We run the same chain in two modes: a real-time engine for preview and an offline engine for the final render.


Real-time preview. We decode about five seconds of audio into memory once, build the engine, and keep it running. When the user taps a preset, we change the node parameters in place and replay the buffer. There is no engine restart, so switching presets feels instant.

func playPreviewChip(_ preset: AudioPreset) {

    player.stop()

    applyPresetToPreviewNodes(preset) // just sets EQ, dynamics, reverb params

    for buf in previewBuffers { player.scheduleBuffer(buf) }

    player.play()

}

Offline export. For the full clip we switch the engine into manual rendering mode. Instead of playing to the speaker, we pull processed audio out 4,096 frames at a time and write it to an AAC file. This runs faster than real time and never touches the audio hardware.

try engine.enableManualRenderingMode(.offline, format: format, maximumFrameCount: 4096)

try engine.start()

playerNode.play()


while engine.manualRenderingSampleTime < totalFrames {

    let frames = AVAudioFrameCount(min(4096, totalFrames - engine.manualRenderingSampleTime))

    let status = try engine.renderOffline(frames, to: renderBuffer)

    if status == .success { try outputFile.write(from: renderBuffer) }

}

The processed audio track is then muxed back with the original video track using a passthrough export, so the video is never re-encoded.


One detail worth knowing: the Dynamics Processor is a classic Audio Unit, not a typed AVAudioUnit subclass. You configure it through AudioUnitSetParameter with kDynamicsProcessorParam_* constants.

What We Got Wrong Along the Way

Denoise ate consonants. The first version gated hard at −28 dB to silence the gaps between words. It worked, but it also swallowed soft word endings like t, p and s, so speech sounded clipped. We relaxed the gate to −32 dB. We also lowered the high-pass from a higher cutoff to 100 Hz, because aggressive low cuts were thinning out male voices, whose chest resonance sits around 85–180 Hz. The lesson: a noise gate tuned on silence must be checked on quiet speech.


Narrow beats wide. Cutting room boxiness around 300 Hz with a wide band made voices sound hollow. A narrow notch (0.25 octaves) at 320 Hz removed the boxiness and left the warmth alone.


A fake visualizer felt wrong. Our preview screen has a glow that reacts to the audio. The first version pulsed on a timer. People noticed immediately that it didn’t match what they heard. We replaced it with a real tap on the mixer that runs an FFT with Accelerate (vDSP), splits it into five roughly logarithmic bands, and smooths it with fast attack and slow release. Now the glow moves with the sound.


The FFT runs on the audio thread, so the analysis function is a nonisolated static method. It cannot accidentally touch main-actor state, and results hop to the main actor only after the math is done.

Key Takeaways

You don’t need a server or a DSP library to ship professional audio processing on iOS. You need a good chain and well-tuned numbers.


  • Build one chain, vary the data. Five Apple audio units in a fixed order cover everything from voice cleanup to cathedral reverb.

  • Share the chain between preview and export. A real-time engine for instant feedback and manual offline rendering for the final render.

  • Change parameters, don’t rebuild the engine. That is what makes preset switching feel instant.

  • Tune on real speech, not on silence. Most of our tuning time went into making sure presets cleaned up noise without hurting the voice.

  • If the UI claims to react to audio, make it react to the actual audio. Users notice when it doesn’t.

MORE BLOGS

Dive Into The
Pixster Blog

Dive Into The
Pixster Blog

Your hub for app updates, design insights, tech trends, and AI innovations.

Stay inspired with creative, cutting-edge content.

Your hub for app updates, design insights, tech trends, and AI innovations. Stay inspired with creative, cutting-edge content.