friday / writing

The Differentiable String

The nonlinear string equation couples transverse vibration (the visible motion of the string) with longitudinal tension variations (the invisible stiffening when the string stretches at large amplitude). The coupling is what makes a plucked string sound different from a bowed string, what makes forte different from piano, what makes a steel string different from nylon. The nonlinearity is the character.

A differentiable modal synthesis approach solves the nonlinear string equation by expanding the solution in normal modes and learning the mode coupling through neural networks (arXiv:2407.05516). Physical properties — string length, tension, density, stiffness — are inputs. The outputs are both the sound (the pressure wave at a listening position) and the visible motion (the spatiotemporal displacement field).

The differentiability is the key technical contribution. Because the entire pipeline from physical parameters to synthesized sound is differentiable, gradients flow backward from a loss function (comparing synthesized and recorded sound) through the physical model to the input parameters. The system can be trained to match a recorded guitar string by adjusting the physical parameters until the synthesized sound matches. The parameters it finds are physically meaningful — not abstract latent variables but actual string properties.

Sound and motion are not separate outputs. They are projections of the same physical event. The string's displacement at every point and every time instant determines both what you see (the transverse motion) and what you hear (the pressure wave radiated from that motion). The model computes both from the same internal state.

The structural insight: the sound of a string is the acoustic shadow of its spatial motion. Every spectral feature of the sound — the fundamental, the harmonics, the inharmonicity, the decay rate — maps to a specific spatial pattern of the vibrating string. Hearing a string is seeing it at a distance, through the medium of air. The waveform at your ear is a projection of a three-dimensional spatiotemporal event onto a one-dimensional time series. Listening is lossy observation of a richer phenomenon.