September 27, 2026
September 27, 2026
How to Check Product Name Pronunciation in AI Voiceovers
Prepare, generate, and review AI ad voiceovers with a reusable sheet for product names, numbers, and pacing.
Prepare, generate, and review AI ad voiceovers with a reusable sheet for product names, numbers, and pacing.
A practical workflow for ecommerce teams to check product names, numbers, pacing, and final audio before using an AI voiceover in an ad.
To check product name pronunciation in an AI voiceover, define the approved spoken form before generation, write numbers as they should be heard, and review the exported audio against a line-by-line sheet. Do not approve a take until a human listener confirms the name, numbers, instructions, and pacing.
This is Kubflow’s suggested production workflow. It uses documented ElevenLabs guidance as an illustrative provider reference, but the review process works with other speech models too.
Prepare the script and pronunciation reference
Start with three approved inputs. This prevents reviewers from debating pronunciation after several voiceovers have already been generated.
Approved product name: Include exact capitalization, accents, and spacing.
Approved spoken form: Write a simple phonetic cue, ideally confirmed by the brand owner or product team.
Commercial details: Confirm prices, quantities, dates, measurements, discount terms, and usage instructions.
For example, a fictional brand might specify that “Nuvéra Glow” must be spoken as “new-VAIR-uh glow.” That spelling is an internal listening reference, not proof that every speech model will pronounce it correctly.
Keep the ad script short while correcting pronunciation. A 15-second script is easier to review than a full edit containing music, effects, and several speakers. Add those elements only after the clean voice track passes.
Check your selected speech model’s official guidance
Pronunciation controls differ by provider and model. Before generating, check the official documentation for the exact speech model you selected rather than assuming that a setting is available.
For example, ElevenLabs documents pronunciation dictionaries that can use IPA or CMU notation for specified terms. It also states that dictionary matching is case-sensitive and that a dictionary version is attached to a particular text-to-speech request. See its pronunciation dictionary documentation.
ElevenLabs also notes that phoneme support varies by model. Its guidance suggests aliases or phonetic spellings where the selected model does not support the relevant phoneme method. These are provider capabilities, not a statement that Kubflow exposes dictionary uploads, aliases, speed controls, normalization, or model parameters.
Use only the controls visibly available in your workflow. If a needed provider control is absent, adjust the script using ordinary wording or prepare the speech through an appropriate external provider interface, then review the resulting audio manually.
Write numbers for the delivery you want
Do not leave commercially important numbers open to interpretation. ElevenLabs advises writing numbers, dates, symbols, and acronyms fully in words when a particular delivery is required. Its documentation also warns that currency, dates, addresses, abbreviations, and similar complex text can be mispronounced.
Review the provider’s guidance for numbers and dates, then make the intended reading explicit in the script.
Write “forty-nine dollars” instead of “$49.”
Write “September thirtieth” instead of “9/30.”
Write “two fifteen-milliliter bottles” if that is the approved reading.
Expand an acronym when customers need to hear the full phrase.
Automatic text normalization may choose a plausible reading, but not necessarily the commercial reading your brand intends. Human review remains necessary.
Run the product name pronunciation AI voiceover workflow
Save the approved script. Use a versioned filename so reviewers know exactly which text produced each take.
Generate a clean voice track. Keep music and sound effects out of the first review export.
Listen without reading. A reviewer should first hear the take as a customer would.
Listen while reading. Compare every product name, number, date, claim, and instruction with the approved script.
Mark the review sheet. Record pass, fail, and a specific reason for every important item.
Revise one problem at a time. Change punctuation, sentence structure, or supported pronunciation controls, then save a new take.
Approve the final audio in context. Place the accepted voiceover under the ad visuals and music, then check that key words remain audible.
Kubflow can connect text, image, video, and audio models on a visual workflow canvas, letting a team connect generation steps and rerun a workflow. The output still needs human review. For a broader review strategy, see what to automate and what to review in AI creative production.
Illustrative script, files, and review decision
This example is hypothetical. It is a reusable recipe, not a customer result or a test performed by Kubflow.
Product: Nuvéra Glow
Approved delivery: new-VAIR-uh glow
File: nuvera_glow_script_v01.txt
Contents: “Meet Nuvéra Glow. Get two bottles for forty-nine dollars, available through September thirtieth. Apply once each evening. Discover Nuvéra Glow today.”
Copyable generation prompt: “Read this ecommerce ad in a warm, clear, conversational voice. Keep the pace steady. Pause briefly after the product name and after the offer. Do not add, remove, or rewrite words. Script: Meet Nuvéra Glow. Get two bottles for forty-nine dollars, available through September thirtieth. Apply once each evening. Discover Nuvéra Glow today.”
Output: nuvera_glow_voiceover_take01.wav
The hypothetical reviewer hears the product name correctly, but “once each” runs together.
File: nuvera_glow_pronunciation_review_v01.csv
Contents:
Item | Intended delivery | Result |
|---|---|---|
Nuvéra Glow | new-VAIR-uh glow | Pass |
Two bottles | Two bottles | Pass |
Forty-nine dollars | Not “four nine dollars” | Pass |
September thirtieth | Full date wording | Pass |
Once each evening | Clear and unhurried | Fail: words run together |
Decision: Reject take 01. Revise the script to “Apply once. Each evening.” and generate nuvera_glow_voiceover_take02.wav. Accept take 02 only after two human reviewers confirm the product name, all numbers, and the revised pacing.
Use this final audio acceptance checklist
The product name matches the brand owner’s approved pronunciation.
Every occurrence of the name is correct, including repeated lines.
Prices, quantities, dates, percentages, and measurements are spoken as intended.
Instructions do not run together or become ambiguous.
The voice has enough space around the offer and call to action.
No words were added, omitted, or changed.
The voice remains understandable under music and sound effects.
The final export matches the approved take and filename.
If the ad is intended for Google Ads, remember that Google requires video-ad information and media to be understandable and to identify the advertised product, service, or entity. Meeting this audio checklist does not guarantee platform approval.
Know when to rewrite instead of regenerate
Regeneration is not always the best fix. Rewrite the line when repeated takes misread the same number, a phonetic spelling sounds unnatural, punctuation creates awkward pauses, or a sentence contains too many commercial details.
Split crowded copy into shorter sentences. If the product name still fails, use a provider-supported pronunciation method where available or record the critical line with a human voice actor. AI generation does not guarantee product-name fidelity, and editorial approval cannot guarantee ad approval or performance.
When your script and review sheet are ready, build the repeatable creative workflow in Kubflow, then keep the human listening gate before final export.
A practical workflow for ecommerce teams to check product names, numbers, pacing, and final audio before using an AI voiceover in an ad.
To check product name pronunciation in an AI voiceover, define the approved spoken form before generation, write numbers as they should be heard, and review the exported audio against a line-by-line sheet. Do not approve a take until a human listener confirms the name, numbers, instructions, and pacing.
This is Kubflow’s suggested production workflow. It uses documented ElevenLabs guidance as an illustrative provider reference, but the review process works with other speech models too.
Prepare the script and pronunciation reference
Start with three approved inputs. This prevents reviewers from debating pronunciation after several voiceovers have already been generated.
Approved product name: Include exact capitalization, accents, and spacing.
Approved spoken form: Write a simple phonetic cue, ideally confirmed by the brand owner or product team.
Commercial details: Confirm prices, quantities, dates, measurements, discount terms, and usage instructions.
For example, a fictional brand might specify that “Nuvéra Glow” must be spoken as “new-VAIR-uh glow.” That spelling is an internal listening reference, not proof that every speech model will pronounce it correctly.
Keep the ad script short while correcting pronunciation. A 15-second script is easier to review than a full edit containing music, effects, and several speakers. Add those elements only after the clean voice track passes.
Check your selected speech model’s official guidance
Pronunciation controls differ by provider and model. Before generating, check the official documentation for the exact speech model you selected rather than assuming that a setting is available.
For example, ElevenLabs documents pronunciation dictionaries that can use IPA or CMU notation for specified terms. It also states that dictionary matching is case-sensitive and that a dictionary version is attached to a particular text-to-speech request. See its pronunciation dictionary documentation.
ElevenLabs also notes that phoneme support varies by model. Its guidance suggests aliases or phonetic spellings where the selected model does not support the relevant phoneme method. These are provider capabilities, not a statement that Kubflow exposes dictionary uploads, aliases, speed controls, normalization, or model parameters.
Use only the controls visibly available in your workflow. If a needed provider control is absent, adjust the script using ordinary wording or prepare the speech through an appropriate external provider interface, then review the resulting audio manually.
Write numbers for the delivery you want
Do not leave commercially important numbers open to interpretation. ElevenLabs advises writing numbers, dates, symbols, and acronyms fully in words when a particular delivery is required. Its documentation also warns that currency, dates, addresses, abbreviations, and similar complex text can be mispronounced.
Review the provider’s guidance for numbers and dates, then make the intended reading explicit in the script.
Write “forty-nine dollars” instead of “$49.”
Write “September thirtieth” instead of “9/30.”
Write “two fifteen-milliliter bottles” if that is the approved reading.
Expand an acronym when customers need to hear the full phrase.
Automatic text normalization may choose a plausible reading, but not necessarily the commercial reading your brand intends. Human review remains necessary.
Run the product name pronunciation AI voiceover workflow
Save the approved script. Use a versioned filename so reviewers know exactly which text produced each take.
Generate a clean voice track. Keep music and sound effects out of the first review export.
Listen without reading. A reviewer should first hear the take as a customer would.
Listen while reading. Compare every product name, number, date, claim, and instruction with the approved script.
Mark the review sheet. Record pass, fail, and a specific reason for every important item.
Revise one problem at a time. Change punctuation, sentence structure, or supported pronunciation controls, then save a new take.
Approve the final audio in context. Place the accepted voiceover under the ad visuals and music, then check that key words remain audible.
Kubflow can connect text, image, video, and audio models on a visual workflow canvas, letting a team connect generation steps and rerun a workflow. The output still needs human review. For a broader review strategy, see what to automate and what to review in AI creative production.
Illustrative script, files, and review decision
This example is hypothetical. It is a reusable recipe, not a customer result or a test performed by Kubflow.
Product: Nuvéra Glow
Approved delivery: new-VAIR-uh glow
File: nuvera_glow_script_v01.txt
Contents: “Meet Nuvéra Glow. Get two bottles for forty-nine dollars, available through September thirtieth. Apply once each evening. Discover Nuvéra Glow today.”
Copyable generation prompt: “Read this ecommerce ad in a warm, clear, conversational voice. Keep the pace steady. Pause briefly after the product name and after the offer. Do not add, remove, or rewrite words. Script: Meet Nuvéra Glow. Get two bottles for forty-nine dollars, available through September thirtieth. Apply once each evening. Discover Nuvéra Glow today.”
Output: nuvera_glow_voiceover_take01.wav
The hypothetical reviewer hears the product name correctly, but “once each” runs together.
File: nuvera_glow_pronunciation_review_v01.csv
Contents:
Item | Intended delivery | Result |
|---|---|---|
Nuvéra Glow | new-VAIR-uh glow | Pass |
Two bottles | Two bottles | Pass |
Forty-nine dollars | Not “four nine dollars” | Pass |
September thirtieth | Full date wording | Pass |
Once each evening | Clear and unhurried | Fail: words run together |
Decision: Reject take 01. Revise the script to “Apply once. Each evening.” and generate nuvera_glow_voiceover_take02.wav. Accept take 02 only after two human reviewers confirm the product name, all numbers, and the revised pacing.
Use this final audio acceptance checklist
The product name matches the brand owner’s approved pronunciation.
Every occurrence of the name is correct, including repeated lines.
Prices, quantities, dates, percentages, and measurements are spoken as intended.
Instructions do not run together or become ambiguous.
The voice has enough space around the offer and call to action.
No words were added, omitted, or changed.
The voice remains understandable under music and sound effects.
The final export matches the approved take and filename.
If the ad is intended for Google Ads, remember that Google requires video-ad information and media to be understandable and to identify the advertised product, service, or entity. Meeting this audio checklist does not guarantee platform approval.
Know when to rewrite instead of regenerate
Regeneration is not always the best fix. Rewrite the line when repeated takes misread the same number, a phonetic spelling sounds unnatural, punctuation creates awkward pauses, or a sentence contains too many commercial details.
Split crowded copy into shorter sentences. If the product name still fails, use a provider-supported pronunciation method where available or record the critical line with a human voice actor. AI generation does not guarantee product-name fidelity, and editorial approval cannot guarantee ad approval or performance.
When your script and review sheet are ready, build the repeatable creative workflow in Kubflow, then keep the human listening gate before final export.








