org.llm4s.speech.tts.provider

Members list

Type members

Classlikes

final class AzureTTSClient(config: TTSConfig, httpClient: Llm4sHttpClient) extends TextToSpeech

Azure AI Speech text-to-speech (REST POST /cognitiveservices/v1, SSML body).

Azure AI Speech text-to-speech (REST POST /cognitiveservices/v1, SSML body).

Requests the raw-24khz-16bit-mono-pcm output format, so GeneratedAudio.data is headerless 24 kHz, 16-bit, mono PCM described by an honest AudioMeta. Write it out with org.llm4s.speech.io.WavFileGenerator.saveAsWav.

options.voice overrides the configured voice name; options.language sets xml:lang (default: the locale prefix of the voice name, else en-US); options.speakingRate becomes a prosody rate.

Value parameters

config

configuration from org.llm4s.speech.config.SpeechConfigLoader.tts; baseUrl is the regional https://<region>.tts.speech.microsoft.com endpoint

httpClient

HTTP transport, replaceable in tests

Attributes

Companion
object
Supertypes
trait TextToSpeech
class Object
trait Matchable
class Any

Attributes

Companion
class
Supertypes
class Object
trait Matchable
class Any
Self type
final class ElevenLabsTTSClient(config: TTSConfig, httpClient: Llm4sHttpClient) extends TextToSpeech

ElevenLabs text-to-speech (POST /v1/text-to-speech/{voice_id}).

ElevenLabs text-to-speech (POST /v1/text-to-speech/{voice_id}).

Requests output_format=pcm_24000 (raw 24 kHz, 16-bit, mono little-endian samples), so GeneratedAudio.data is headerless PCM described by an honest AudioMeta. Write it out with org.llm4s.speech.io.WavFileGenerator.saveAsWav. options.voice overrides the configured voice id.

Value parameters

config

configuration from org.llm4s.speech.config.SpeechConfigLoader.tts; model is the ElevenLabs model_id and voice the voice id

httpClient

HTTP transport, replaceable in tests

Attributes

Companion
object
Supertypes
trait TextToSpeech
class Object
trait Matchable
class Any

Attributes

Companion
class
Supertypes
class Object
trait Matchable
class Any
Self type
final class OpenAITTSClient(config: TTSConfig, httpClient: Llm4sHttpClient) extends TextToSpeech

OpenAI text-to-speech (POST /v1/audio/speech).

OpenAI text-to-speech (POST /v1/audio/speech).

Asks for response_format = pcm, which OpenAI documents as raw 24 kHz, 16-bit, mono little-endian samples, so GeneratedAudio.data is headerless PCM described by an honest AudioMeta - the same shape org.llm4s.speech.tts.Tacotron2TextToSpeech produces. Write it out with org.llm4s.speech.io.WavFileGenerator.saveAsWav.

options.voice overrides the configured voice; options.speakingRate is sent as speed (OpenAI accepts 0.25 to 4.0).

Value parameters

config

configuration from org.llm4s.speech.config.SpeechConfigLoader.tts

httpClient

HTTP transport, replaceable in tests

Attributes

Companion
object
Supertypes
trait TextToSpeech
class Object
trait Matchable
class Any

Attributes

Companion
class
Supertypes
class Object
trait Matchable
class Any
Self type