How to Build an Audiobook Pronunciation Guide
Create a usable pronunciation guide for names, places, technical terms, invented words, and multilingual passages without slowing production.

A pronunciation guide is a production tool, not a decorative glossary. Its job is to let a narrator or voice system make the same defensible choice every time a difficult word appears. A list of spellings without audio, context, or a decision owner does not do that job.
Build the guide before full narration, but expect it to grow. The manuscript will reveal edge cases during recording: a surname that changes after an identity reveal, a scientific abbreviation spoken differently in dialogue, or a place name whose local pronunciation differs from the obvious reading.
Extract candidates systematically
Begin with a scan of the full manuscript. Collect:
- character names, surnames, titles, and aliases;
- real and invented place names;
- organizations, brands, and product names;
- foreign-language words and phrases;
- historical figures and cultural references;
- scientific, medical, legal, and technical terms;
- abbreviations, initialisms, formulas, and units;
- invented species, magic systems, and world-building vocabulary;
- words whose pronunciation changes by meaning.
Do not include every ordinary word. A guide becomes hard to use when uncertain items are buried among hundreds of obvious entries. Mark each candidate with chapter, sentence, speaker, and meaning. Context distinguishes “lead” the metal from “lead” the verb and tells you whether “Dr.” should be spoken as “Doctor.”
Search is useful but not sufficient. Capitalization can locate proper nouns, yet it also finds words at the start of sentences and misses lower-case technical terms. Combine automated extraction with an editorial pass.
Decide which authority applies
Different words require different sources. For a living person's name, that person's own recorded introduction is stronger evidence than a generic dictionary. For a place, local government, public broadcaster, university, or transport announcement may establish local usage. For a technical term, a field-specific dictionary or the organization that coined it can be useful.
Use a source hierarchy:
- the author's explicit direction for fictional terms;
- the named person's own speech or an official source;
- an authoritative field or language reference;
- several credible recordings from the relevant region;
- a documented editorial decision when no authority settles the case.
Record the source link and access date. Web pages change, and a producer should be able to explain why an unusual reading was chosen. Do not cite an anonymous pronunciation video as if it were conclusive simply because it ranks first.
For another language, ask a qualified speaker about the exact word in context. Pronunciation can change with grammar, neighboring sounds, dialect, or the speaker's relationship to the language. The goal is not to flatten every word into an English approximation, nor to switch suddenly into a performance that breaks the character.
Store something people can actually perform
An entry should contain more than phonetic spelling:
| Field | Purpose |
|---|---|
| Display form | Exact manuscript spelling |
| Spoken form | Plain-language cue or syllable breakdown |
| IPA | Precise representation when the performer can use it |
| Stress | The syllable that carries emphasis |
| Audio reference | Fastest way to hear the target |
| Context | Meaning, speaker, and chapter |
| Source | Evidence and access date |
| Status | proposed, author-approved, or locked |
| Notes | allowed variants and character-specific treatment |
Plain-language respelling is accessible but ambiguous. “MAY-uh” may still leave the vowel quality uncertain. IPA is precise but only useful to someone who can read it. A short reference recording plus a written cue serves the widest range of collaborators.
Keep audio clips short, normalized to a comfortable level, and named with a stable identifier. The clip should say the term alone and in a short phrase. Context exposes stress changes that an isolated word hides.
Ask the author focused questions
Sending an author a spreadsheet of two hundred blank pronunciation cells creates delay. Research first, group questions, and identify the few decisions only the author can make.
Use questions such as:
- “Is Aveline pronounced AV-uh-line or AV-uh-leen?”
- “Should the initials N.C.R. be spoken letter by letter?”
- “This town has two documented local variants; which fits the character?”
- “Does the disguised character keep the same accent after the reveal?”
Avoid asking “How do I pronounce everything?” Provide evidence and a proposed answer. Mark approval in the guide rather than relying on a message thread that will be difficult to find during chapter twenty.
Handle consistency without erasing character
The same word does not always need one universal performance. A child may mispronounce a name intentionally. A newcomer may use an outsider's version of a place name. A character may code-switch. These are story facts, not errors, when the manuscript supports them.
Represent intentional variants as separate entries tied to a speaker or scene. The canonical production reading remains visible, and the exception includes a reason. This prevents quality control from “fixing” an intentional mistake.
Conversely, do not invent variants to make voices more distinctive. Character contrast should come from the voice casting system, not random changes to proper nouns.
Integrate the guide into rendering
A guide stored in a forgotten document does not protect the audio. Put pronunciation checks at three moments:
- Before rendering a chapter: surface every guide entry that appears in the chapter.
- During spot review: compare the first occurrence and any high-risk occurrence against the reference.
- During final QC: search the audio defect log for pronunciation changes and confirm corrections did not create new inconsistencies.
For a human narrator, provide the chapter subset rather than forcing constant searching through the master list. For a speech system, use a supported pronunciation dictionary, phoneme control, or carefully tested text substitution. Never alter the canonical manuscript just to force output; keep a separate render representation and preserve the source.
Text substitutions need boundaries. Replacing read globally will damage both
present and past tense. Replacing a short name inside longer words can create
new errors. Match complete tokens where possible, attach context, and test the
rendered sentence.
Review abbreviations and numbers separately
Pronunciation work includes decisions that are not traditional dictionary entries:
- Does
2026mean “twenty twenty-six” or “two thousand twenty-six”? - Is
NASAa word whileFBIis spoken as letters? - Does
St.mean Saint or Street here? - Should
3 × 4be “three by four” or “three times four”? - How are URLs, handles, citations, and footnote markers treated?
Create project-level rules for recurring forms, then record exceptions. Units may expand differently in narrative prose and dialogue. A character can say “five K” while the narrator reads “five kilometers” if the text and direction support that difference.
Run a pronunciation lock
Before the long render, sort all entries by status. Resolve every high-frequency or story-critical item. Low-risk unresolved terms can remain flagged, but they need an owner and deadline.
Listen to a test scene containing several difficult terms. Check that the pronunciation is correct, natural at sentence speed, and compatible with the character voice. A technically accurate phoneme sequence can still sound awkward when stress or surrounding pauses are wrong.
After approval, lock the entry. If later evidence changes the decision, log the new version and identify every completed chapter containing the old form. That makes correction scope knowable.
A compact guide beats repeated repair
The cost of a pronunciation error is not one bad word. It is finding every occurrence, recreating the surrounding performance, matching the audio, and re-running quality control. A guide moves that work to the cheaper side of production.
In VoxForge, you can test difficult lines with the assigned character voice before rendering a full chapter. Keep the approved pronunciation and context in your production ledger, then compare the first rendered occurrence. The tool provides repeatability; the guide supplies editorial judgment.
Extract with context, research from the right authority, save audio references, and track approval. That is enough to turn pronunciation from recurring guesswork into a controlled part of audiobook production.