Skip to content

Sources and method

This is the page that makes the rest of the site checkable. It states exactly where the technical content came from, exactly what was excluded and why, and gives a link to every source so you can go and disagree with us.

Method

The technical material here comes from four kinds of source:

  1. US patents. Ten of them, all assigned to Digital Voice Systems, Inc., granted between 1992 and 1999. A patent must enable the invention it claims, which makes it an unusually complete public teaching document. These are the backbone of the site.
  2. Published academic literature. The Griffin and Lim 1988 paper that introduced the Multi-Band Excitation model, and the Hardwick and Lim 1988 paper that turned it into a 4.8 kbps coder. Both are in the normal scholarly record.
  3. Public system specifications. Principally the JARL D-STAR system specification, which defines the over-the-air frame structure, plus publicly available DVSI product documentation for device-level behaviour.
  4. Our own measurements of commercial hardware. A DVSI AMBE-3000 device operated as a black box: audio and frames in one side, recorded output from the other.

Claims name their source inline, in the form US 5,701,390 or Griffin & Lim 1988, so you can go and read the original. Inferences and engineering judgement are marked as such.

How the hardware measurements were made

The one class of source here that isn't a document is our own bench work, so it deserves its own method statement.

The device is a ThumbDV-class USB vocoder stick containing a DVSI AMBE-3000 chip, bought at retail. It is driven over its documented serial protocol, using the packet format described in DVSI's own publicly available product documentation. Audio or encoded frames go in; encoded frames or audio come out; both sides are recorded.


Bibliography

Notes flag the links that need a browser rather than a command-line fetch.

US patents, the primary technical sources

All assigned to Digital Voice Systems, Inc. Expiry dates and legal-status detail are on The patent landscape.

Patent Title Used for
US 5,081,681 Method and apparatus for phase synthesis for speech processing Coherent-plus-jittered phase synthesis
US 5,216,747 Voiced/unvoiced estimation of an acoustic signal Pitch tracking, energy-adaptive voicing thresholds
US 5,247,579 Methods for speech transmission Adaptive spectral enhancement, frame repeat, error-rate smoothing
US 5,630,011 Quantization of harmonic amplitudes representing speech Predictive residual-DCT amplitude quantization
US 5,649,050 Maintaining data rate integrity despite mismatch of readiness between components Buffering and rate adaptation around a vocoder
US 5,701,390 Synthesis of MBE-based coded speech using regenerated phase information Regenerated-phase synthesis, frame-boundary rules
US 5,715,365 Estimation of excitation parameters Nonlinear pre-processing for pitch and voicing
US 5,754,974 Spectral magnitude representation for multi-band excitation speech coders Voicing-independent magnitudes, unvoiced synthesis, analysis window
US 5,826,222 Estimation of excitation parameters (continuation) Hybrid pitch estimator, voicing smoothing
US 5,870,405 Digital transmission of acoustic signals over a noisy communication channel FEC, bit prioritization, scrambling, frame repeat and mute

Later DVSI patents referenced for landscape purposes only, none of them a source of technical content here: US 6,161,089 · US 6,199,037 · US 6,912,495 · US 7,634,399 · US 7,957,963 · US 7,970,606 · US 8,036,886 · US 8,200,497 · US 8,315,860 · US 8,359,197 · US 8,595,002.

Papers

D. W. Griffin and J. S. Lim, "Multiband Excitation Vocoder," *IEEE
Transactions on Acoustics, Speech, and Signal Processing*, vol. 36, no. 8,
pp. 1223–1235, August 1988.
doi:10.1109/29.1651
· scanned PDF, qsl.net

The paper that defines the model: speech as a spectral envelope times an excitation spectrum, with an independent voiced/unvoiced decision per frequency band. Source for the model itself, the analysis-by-synthesis error criterion, and the shape of both synthesizers. The DOI link goes to IEEE Xplore, which returns a bot challenge to command-line clients but opens normally in a browser; the qsl.net scan is a full copy of the same paper.

**J. C. Hardwick and J. S. Lim, "A 4.8 kbps multi-band excitation speech
coder,"** ICASSP-88, pp. 374–377, 1988.
doi:10.1109/ICASSP.1988.196595
· full text, archive.org
· MIT thesis of the same title, DSpace@MIT

The step from model to codec — quantization and bit allocation for a real 4.8 kbps system. Hardwick is the named inventor on most of the DVSI patents above, and this is the direct ancestor of IMBE. Same IEEE bot-challenge caveat on the DOI link.

D. W. Griffin and J. S. Lim, "Multi-Band Excitation Vocoder," Rome Air
Development Center technical report, DTIC
ADA181146.

The MBE work was funded in part by the US Air Force (RADC, Griffiss AFB, contract F19628-85-K-0028), which makes the contract deliverables public technical reports rather than proprietary documents. Contains DRT intelligibility scores for clean and noisy speech.

Specifications

JARL D-STAR system specification.
https://www.jarl.com/d-star/shogen.pdf

The Japan Amateur Radio League's published specification for D-STAR. Source for everything on this site about the over-the-air layer: the 4800 bps GMSK channel, the 20 ms frame, and the split of that frame into voice and data. The URL serves an English-language PDF.

TIA-102.BABA-1, "APCO Project 25 Half-Rate Vocoder Addendum." In the
TIA-102 series document collection on archive.org.

A public specification of the 3600 bps half-rate vocoder, and the document the amateur community generally equates with AMBE+2 half-rate. Freely readable, and a legitimate source for anyone working on that codec. This site draws little from it for the simple reason that its subject is full-rate AMBE as used by D-STAR. See The patent landscape for the half-rate patent position.

Vendor documentation

Digital Voice Systems, Inc. publishes product-level documentation for its vocoder chips and software. It is used here only for device-level and protocol-level facts — packet formats, rate options, product capability claims.

Background and context

These informed the site's framing and its sense of what readers already believe. None of them is a technical source; where any of them disagrees with a patent or a specification, the primary document governs.

Verification tools

  • Google Patents — full text, claims and legal-status timelines. Convenient; its computed dates and its OCR of equations both need checking.
  • USPTO Patent Center — the authoritative file wrapper: term adjustment, maintenance fees, terminal disclaimers. Governs where the two disagree.

Licensing

Written by Rob Ludwick, AJ7HR.

The prose, diagrams and page content are licensed CC BY 4.0. The tooling, build scripts, stylesheets, animation runtime, CI workflows is MIT. Reuse the writing with attribution. Reuse the tooling under MIT.

The speech on Listen was synthesised with Piper, not recorded. The male clips use a voice trained on LibriTTS-R (OpenSLR 141), which is CC BY 4.0 and requires the credit speech synthesised with a Piper voice trained on LibriTTS-R (OpenSLR 141), used under CC BY 4.0. The corpus is Koizumi et al., LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus, 2023. The female clips use a voice trained on the LJ Speech Dataset, which is public domain. docs/assets/audio/MANIFEST.md records both, the model checksums, and a fine-tuning lineage on the male voice that it states rather than resolves.

AMBE, AMBE+, and AMBE+2 are trademarks of Digital Voice Systems, Inc., used here only to identify the technology under discussion. This project is independent of DVSI and is not affiliated with, sponsored by, endorsed by, or approved by them.