Speech production and speech perception are opposite processes. Production converts intention into motor commands — the brain plans articulation, activates muscles, generates sound. Perception converts sound into meaning — the ear receives pressure waves, the auditory cortex processes them, language areas extract semantics. One is output; the other is input. The neural pathways are different. The timescales are different. The computational demands are different.
The paper (arXiv:2603.12628, March 2026) builds a single neural decoder that works for both. The same model, trained on brain activity during both speech production and speech perception, translates neural signals to text in both directions. The decoder doesn't need separate modules for hearing and speaking — a unified representation handles both.
The implication: the brain uses shared representations for processes that appear computationally opposite. At some level of neural encoding, hearing a word and saying a word activate overlapping patterns. The decoder discovers this overlap by being forced to handle both tasks with the same parameters. It works because the overlap is real — not a training artifact but a reflection of how language is organized in cortex.
This connects to the motor theory of speech perception — the hypothesis that perceiving speech involves simulating the motor commands that would produce it. The shared decoder provides indirect evidence: if the neural codes for production and perception were truly independent, no single decoder could handle both. That it works suggests the codes share structure.
The structural lesson: two processes that seem opposite may share a representation that is neither input nor output but something more abstract — a language-level encoding that both production and perception access. The opposition (input vs output) is a property of the interface with the world, not of the internal representation. Inside the brain, speaking and listening may be the same computation viewed from different sides.