Aristech provides the SSML documentation for each language via the API in JSON format.

The documentation for the voice anne_de_DE_v8 is presented below as a human-readable overview.

Overview of Supported Tags

<audio><break><lang><phoneme><prosody><voice>
✅✅❌✅✅✅

<audio> — Audio-Integration

Embed audio files (e.g., jingles, sound effects, music) directly into the synthesized speech output. The audio files must be in WAV format and available in the server's configured audio directory (the .wav extension is optional). The audio is automatically adjusted to the target sample rate.

<audio name="jingle"/> Welcome to our service.
<audio name="fanfare"/> <break time="600ms"/> The winner will be announced now.

<break> — Pauses

Inserts a pause of a specified duration at any point. The time attribute specifies the duration of the pause in milliseconds (ms) or seconds (s). Valid range: 1 ms to 120 seconds (typically 100–600 ms).

One moment please <break time="200ms"/> this is good.
A longer pause: <break time="1.5s"/> then let's continue.
<break time="50ms"/>

Punctuation marks at the end of a sentence (., !, ?) automatically create appropriate pauses, so explicit break tags between sentences are usually not necessary.

<phoneme> — Pronunciation Control

Specifies the exact phonetic pronunciation of a word or phrase. The complete sequence of sounds must be specified in the ph attribute.

We will say <phoneme ph="j oo0 h a1 n">Johann</phoneme> like this.

Use cases: Proper nouns, foreign words, technical terms, and correcting automatic pronunciation.

<prosody> — Speaking Speed and Volume

Customizable and nestable.

Note: Pitch control is not supported by this voice. Only rate and volume have an effect.

Rate

Standard value 1.0, recommended values 0.8–1.3.

<prosody rate="1.2">This will be spoken faster.</prosody>
<prosody rate="0.8">This will be spoken slower.</prosody>

Volume

Relative volume. Standard value 1.0, recommended range 1.0–3.0. Values which are too high will result in distortion (clipping).

<prosody volume="1.5">This is louder.</prosody>

<voice> — Change of Voice

Switches to another installed voice within a single synthesis request. The name attribute must be a valid voice ID.

<voice name="tom_de_DE">Hello, here is Tom.</voice>
<voice name="anne_de_DE">And I am anne.</voice>

Use cases: dialogue systems, multilingual content, distinguishing between the narrator and a character.

Note: <lang> is not supported for this voice.

  • No labels