AzureSTTClient

org.llm4s.speech.stt.provider.AzureSTTClient
See theAzureSTTClient companion object
final class AzureSTTClient(config: STTConfig, httpClient: Llm4sHttpClient) extends SpeechToText

Azure AI Speech speech-to-text (short-audio REST API, POST /speech/recognition/conversation/cognitiveservices/v1).

The REST endpoint recognises at most about 60 seconds of audio and takes WAV (PCM) input: FileAudio, BytesAudio and StreamAudio are all read as WAV data, and the sample rate is read from the WAV header, falling back to the input's sampleRate, then 16 kHz.

The recognition language is options.language, else the configured default (model of the config, e.g. azure/en-US).

Value parameters

config

configuration from org.llm4s.speech.config.SpeechConfigLoader.stt; baseUrl is the regional https://<region>.stt.speech.microsoft.com endpoint

httpClient

HTTP transport, replaceable in tests

Attributes

Companion
object
Graph
Supertypes
trait SpeechToText
class Object
trait Matchable
class Any

Members list

Value members

Concrete methods

override def transcribe(input: AudioInput, options: STTOptions): Result[Transcription]

Transcribe audio to text.

Transcribe audio to text.

Value parameters

input

Audio data to transcribe

options

Configuration for transcription

Attributes

Returns

Result containing Transcription or STTError

Definition Classes

Inherited methods

def isAvailable: Result[Boolean]

Check if this provider is available/healthy. Useful for failover logic and availability checks.

Check if this provider is available/healthy. Useful for failover logic and availability checks.

Attributes

Inherited from:
SpeechToText

Concrete fields

override val name: String

Unique identifier/name of this provider

Unique identifier/name of this provider

Attributes

override val supportedFormats: List[String]

List supported audio formats (e.g., "audio/wav", "audio/mp3")

List supported audio formats (e.g., "audio/wav", "audio/mp3")

Attributes