Constantin LabsDossier IV of IVForward Pro · EU  +  Four Checkmen · USSpeech
Measured results · known ground truth

Speaker identity is one dimension. Everything else is content.

A representation of speech in which the speaker occupies exactly one dimension and every other dimension carries phonetic content. Not an approximation of that separation — the exact separation, proven maximal and confirmed by measurement, beating the standard method on both axes at once.

one frame of speech speaker phonetic content
one dimension carries the speaker
every other dimension is content
Terms of access

What is shown in this dossier is a presentation of measured results. The technology itself is transferred only under an agreed technological cooperation, accepted by all parties involved. What can be made available under such an agreement extends beyond what is presented here.

The foundation

Why the speaker is exactly one dimension

The representation carries a proof of maximality. For a configuration of a given size we can state the exact number of independent invariants that survive speaker normalisation — not a bound and not an estimate — and show that this construction attains that number.

speaker: 1 dimension  ·  phonetic content: every remaining dimension

The standard alternative loses two dimensions of phonetic information in the course of removing the speaker. This construction loses none, and has nothing to tune.

NormalisationPhonetic dimensions keptLost
Standard approachall but three2
Canonical normalisationall but one0

Speaker variation acts along exactly one direction of the representation. The normalisation removes that direction and nothing else, which is why no phonetic information is lost in the process. It is computed in closed form, on every frame, with nothing to estimate and nothing to tune. The construction and its proof are held as trade secret and are communicated within a cooperation.

Result 1 · speaker normalisation

It beats the standard on both axes at once

Six vowels, five speakers with vocal-tract lengths between 0.85× and 1.25×, with unbalanced vowel distributions — the realistic case. VTLN estimates the correction from the speaker mean, which is contaminated by which vowels were actually uttered. The canonical normalisation is computed on every frame and needs nothing at all.

Representation spacesComparison against VTLN
RepresentationVowelSpeakerUseful / parasitic
Raw formants0.8690.8691.00
Classical VTLN — requires speaker identity0.8740.8111.08
Canonical normalisation — requires nothing0.8920.3872.30

222 examples · chance level: 0.167 vowel, 0.200 speaker

+2.1%
phonetics vs VTLN
−52.3%
speaker vs VTLN
+114%
useful / parasitic ratio
Result 2 · precision

Two orders of magnitude above the industry standard

The estimator operates in a regime where the model applies exactly rather than approximately. That is the whole of the difference. The industry standard fits an approximation everywhere; this fits an exact description where an exact description exists.

Formant precision
MethodF₁F₂F₃F₄
LPC-16 — industry standard8.8212.872.0637.26
Direct estimation0.0260.0140.50.088
  error on bandwidths0.050.030.711.24

error in Hz · ground truth: 600/1400/2500/3400 Hz, bandwidths 80/110/140/180 Hz

0.026
Hz error on F₁
0.05
Hz error on the F₁ bandwidth
340×
more precise than LPC
Result 3 · compression

Less distortion at the same parameter budget

The same measurement, the same signal, the same number of parameters. At 20 parameters: 5.10 dB against 7.25 dB. At 40 parameters transparency is reached, which LPC does not reach at any order.

Coding efficiency
Parameters per frameCanonical representationLPC
1210.02 dB
166.43 dB
205.10 dB7.25 dB
322.62 dB
400.77 dB — transparentnot reached
Result 4 · bit rate

1.897 kbps — 135 times less than PCM

A regularity result bounds how fast this representation can change, and therefore how often it has to be sampled at all. Standard practice samples close to seven times more often than necessary. Every additional frame is interpolable from its neighbours: it carries no information, but it costs bits.

Bit rate
1.897
kbps total
135×
vs PCM
33.7×
vs G.711 telephony
1.97
dB — transparent

Each component has its own physiological rate

A result that emerged from the data and was not specified in the theory: the different parts of the speech apparatus move at different speeds, and each component of the representation inherits a rate of its own. Allocating one rate to all of them wastes bits on the slow components and starves the fast ones — which is what every fixed-frame codec does.

Componentkbps
Spectral envelope1.008
Pitch0.275
Voicing amplitudes0.267
Glottal parameter0.347
Total — transparent at LSD 1.97 dB1.897
5.62 MB → 42.7 KB
a 3-minute call
230 → 1.71 TB
1,000 agents, one year
99.26%
storage saved

Prediction and measurement compose. The theory predicted a rate per frame on the assumption that frames are independent, and measurement confirmed that prediction inside its own assumption. Measurement then found something the prediction could not contain, precisely because of that assumption: the number of frames actually required is nearly seven times smaller. Theory and experiment did not merely agree — the experiment extended the theory.

Result 5 · control

Every parameter is a physical object

Move the second formant and the vowel changes. In a neural codec there is no parameter you can reach for.

Editability

The same property makes voice conversion, accent correction and controlled synthesis possible — operations on quantities that carry acoustic meaning.

Fields of application

Where each result applies

Conditions of validity

  • Maximum precision is reached on the low register. Across the rest of the register the advantage over the industry standard remains two orders of magnitude.
  • Result 1 uses uniformly scaled tracts, the standard model of between-speaker variation.
  • Transparency is guaranteed for the spectral envelope at the stated rate. Fine harmonic structure is regenerated rather than transmitted, which is what keeps the rate where it is.

Immediate extensions

  • Biometric identification — the speaker parameter is isolated and correlates with physiology; error rates have not yet been measured.
  • Vocal state — the relevant glottal parameters and jitter are extracted and validated; the link to affective states requires a labelled corpus.
  • Audible synthesis — the components exist, the signal has not yet been produced.
← Back to Constantin Labs    I · Parametric operators    II · Parametric pulse compilation    III · The physics learning certificate

This dossier is a presentation of measured results and an invitation to establish contact. It is not an offer or solicitation to buy or sell any security, not investment advice, and not a commitment to contract. Access to the technology is granted only under an agreed technological cooperation accepted by all parties involved. Figures come from controlled experiments against known ground truth; conditions of validity are stated within each dossier. Third-party names appear as factual references to publicly known products and organizations and do not indicate any endorsement, affiliation or partnership.

© 2026 Constantin Labs. All rights reserved.