Procedural sound design¶
- Status: Implemented; bedside loudness and acoustic synchronization still need final tuning on the XT1058
- Format: deterministic 16-bit mono PCM at 22,050 Hz
The clock ships no copied movie sound and no stock sample pack. Its buzz,
brownout, crackle, release tick, and landing clack are synthesized in Java,
written as short private-cache WAV files, loaded into one SoundPool, and
deleted after loading.
This is both a licensing choice and an engineering one: a sound can be driven by the exact visual envelope that the user sees.
Signal path¶
flowchart LR
Planned[Planned visual effect] --> Envelope[Deterministic brightness envelope]
Envelope --> Synth[SoundSynth on worker thread]
Synth --> PCM[16-bit mono PCM]
PCM --> WAV[Temporary WAV]
WAV --> Pool[SoundPool]
Pool --> Policy{Sound policy allows play?}
Policy -->|yes| Speaker[System sound stream]
Policy -->|no| Mute[No queued playback]
Pool --> Delete[Delete temporary file after load]
Nixie effects are synthesized while they are still planned, usually many seconds before display. Playback begins first and the visual waits an estimated 60 ms audio lead. That estimate comes from the API-22 mixer configuration and still needs acoustic tuning by ear; Android 5.1 screen recording contains no audio.
Why an ordinary low hum failed¶
The first implementation concentrated energy at 50–360 Hz. The XT1058 speaker reproduces very little below roughly 400 Hz, so a plausible waveform was nearly inaudible on the actual device.
The revised design uses the missing-fundamental effect. A harmonic series can be perceived as having a 120 Hz pitch even when the first few partials are too weak for the speaker. Energy moves into the speaker's useful midrange while the subjective electrical pitch stays low.
Sampling and interpolation¶
With sample rate \(f_s=22050\), sample \(n\) corresponds to
Visual envelopes accept integer milliseconds, so audio samples linearly interpolate adjacent visual samples. With \(j=\lfloor t_n\rfloor\) and \(q=t_n-j\),
That keeps an audio clip smooth without creating a second, subtly different model of the visual event.
Harmonic electrical buzz¶
The shared hum oscillator has fundamental \(f_0=120\) Hz and at most 25 partials. A partial is omitted if it reaches Nyquist. Its normalized sample is
with seeded phases \(\phi_k\) and amplitudes
The low partials remain present but deliberately faint. Partials 4–25 place most energy between 400 Hz and 3 kHz. Accumulated oscillator phase is wrapped to \([0,2\pi)\) so high partials retain float precision over the clip.
Buzz¶
For visual brightness \(b(t)>1\), the audio control level is
The result is \(\ell(t)\) times the harmonic oscillator plus a small band-shaped noise term. A 5 ms raised-cosine window at both ends prevents a hard edge:
A flat visual envelope therefore produces digital silence rather than a constant bedroom drone.
Brownout¶
The visual sag controls depth
and bends the inferred fundamental down by as much as 25%:
Depth scales the harmonic buzz and a small noise term. At the start, a short low-pass noise burst and a damped 380 Hz sinusoid create a midrange clunk:
The time constants use milliseconds in the implementation. A first-order DC blocker then keeps the clip close to zero mean.
Stutter¶
The synthesizer finds each visual interval below 0.97, locates its minimum, and emits one 18–30 ms decaying noise burst there. Five to nine seeded impulses inside the burst add dry contact clicks. Dips closer than 20 ms are coalesced.
This is intentionally broadband; unlike the original low-frequency sounds, the crackle survives the phone speaker. A regression test protects the first dip explicitly—the original sentinel arithmetic could overflow and suppress every stutter sound.
Hum¶
The slow visual hum is silent. Continuous audio would change the clock from a rare characterful object into a source of room noise.
Split-flap mechanism¶
Each flap clip contains three components:
- a six-millisecond high-passed noise tick at release;
- two or three faint random impulses between 40 ms and shortly before landing; and
- a 90 ms landing clack: quickly decaying low-pass noise plus a damped, seeded 700–1000 Hz sinusoid.
Landing begins at the renderer's FlipTimeline.FALL_MILLIS, currently 400 ms,
so audio and mechanics share one timing constant. Four variants are prepared
when sounds are enabled. Playback chooses a variant with ±4% rate and ±10%
volume; simultaneous hour and minute flips offset the second card by 20–40 ms
instead of producing a perfectly doubled sample.
Filtering and output level¶
The one-pole high-pass used for DC blocking and transient shaping is
where \(p\) is 0.995 for the brownout clip and 0.9 for the stutter clip. The release tick uses the same form with 0.7.
After synthesis, a clip with floating peak \(P\) is normalized to no more than 60% of signed 16-bit full scale. For intensity \(i\), electrical sounds use
Flap clips use \(g=1\). Playback volume is then bounded again by the app's stored volume and cap and by Android's system-sound stream level.
Spectral regression tests¶
The tests do more than compare array lengths. A direct discrete Fourier transform over a representative window calculates bin energy
The protected ratio is
Buzz and brownout fixtures require \(R_{\ge400}\ge0.70\); stutter requires at least 0.80. Additional tests protect clip duration, determinism by seed, variation across seeds, silence for flat envelopes, transient placement, intensity ordering, the 0.6 peak limit, and absence of unintended energy from partials above the chosen 5 kHz band.
Playback policy and lifecycle¶
Sound is decorative, off by default, and never communicates unique state. A play is rejected, in order, when:
- sounds are disabled;
- motion is disabled;
- the ringer is silent or vibrate;
- Bluetooth A2DP is connected;
- another app is playing music; or
- quiet hours are active and the session-only preview exception is not.
The SoundPool uses USAGE_ASSISTANCE_SONIFICATION and does not request audio
focus, so the clock never ducks or pauses music. It is created on resume and
released on pause. The private clock-sounds cache is cleared at start and
stop, and no clip queues across a lifecycle boundary.
The last validation step is physical rather than mathematical: listen at bedside distance, compare the audio lead with a camera recording, test silent/vibrate and Bluetooth mute behavior, and tune the conservative default volume on the real XT1058 speaker.