org.llm4s.speech.stt.provider
Members list
Type members
Classlikes
Azure AI Speech speech-to-text (short-audio REST API, POST /speech/recognition/conversation/cognitiveservices/v1).
Azure AI Speech speech-to-text (short-audio REST API, POST /speech/recognition/conversation/cognitiveservices/v1).
The REST endpoint recognises at most about 60 seconds of audio and takes WAV (PCM) input: FileAudio, BytesAudio and StreamAudio are all read as WAV data, and the sample rate is read from the WAV header, falling back to the input's sampleRate, then 16 kHz.
The recognition language is options.language, else the configured default (model of the config, e.g. azure/en-US).
Value parameters
- config
-
configuration from org.llm4s.speech.config.SpeechConfigLoader.stt;
baseUrlis the regionalhttps://<region>.stt.speech.microsoft.comendpoint - httpClient
-
HTTP transport, replaceable in tests
Attributes
- Companion
- object
- Supertypes
Attributes
- Companion
- class
- Supertypes
-
class Objecttrait Matchableclass Any
- Self type
-
AzureSTTClient.type
OpenAI speech-to-text (POST /v1/audio/transcriptions, the Whisper API).
OpenAI speech-to-text (POST /v1/audio/transcriptions, the Whisper API).
FileAudio is uploaded as it is; BytesAudio and StreamAudio are WAV data, as for the other STT engines, and are staged in a temporary file that is deleted afterwards.
options.language is sent as the ISO 639-1 prefix of the tag (en-US becomes en), options.prompt as prompt, and options.enableTimestamps switches the request to verbose_json with word granularity.
Value parameters
- config
-
configuration from org.llm4s.speech.config.SpeechConfigLoader.stt
- httpClient
-
HTTP transport, replaceable in tests
Attributes
- Companion
- object
- Supertypes
Attributes
- Companion
- class
- Supertypes
-
class Objecttrait Matchableclass Any
- Self type
-
OpenAISTTClient.type