org.llm4s.speech.stt.provider.AzureSTTClient
See theAzureSTTClient companion object
final class AzureSTTClient(config: STTConfig, httpClient: Llm4sHttpClient) extends SpeechToText
Azure AI Speech speech-to-text (short-audio REST API, POST /speech/recognition/conversation/cognitiveservices/v1).
The REST endpoint recognises at most about 60 seconds of audio and takes WAV (PCM) input: FileAudio, BytesAudio and StreamAudio are all read as WAV data, and the sample rate is read from the WAV header, falling back to the input's sampleRate, then 16 kHz.
The recognition language is options.language, else the configured default (model of the config, e.g. azure/en-US).
Value parameters
- config
-
configuration from org.llm4s.speech.config.SpeechConfigLoader.stt;
baseUrlis the regionalhttps://<region>.stt.speech.microsoft.comendpoint - httpClient
-
HTTP transport, replaceable in tests
Attributes
- Companion
- object
- Graph
-
- Supertypes
Members list
In this article