OpenAISTTClient

org.llm4s.speech.stt.provider.OpenAISTTClient
See theOpenAISTTClient companion object
final class OpenAISTTClient(config: STTConfig, httpClient: Llm4sHttpClient) extends SpeechToText

OpenAI speech-to-text (POST /v1/audio/transcriptions, the Whisper API).

FileAudio is uploaded as it is; BytesAudio and StreamAudio are WAV data, as for the other STT engines, and are staged in a temporary file that is deleted afterwards.

options.language is sent as the ISO 639-1 prefix of the tag (en-US becomes en), options.prompt as prompt, and options.enableTimestamps switches the request to verbose_json with word granularity.

Value parameters

config

configuration from org.llm4s.speech.config.SpeechConfigLoader.stt

httpClient

HTTP transport, replaceable in tests

Attributes

Companion
object
Graph
Supertypes
trait SpeechToText
class Object
trait Matchable
class Any

Members list

Value members

Concrete methods

override def transcribe(input: AudioInput, options: STTOptions): Result[Transcription]

Transcribe audio to text.

Transcribe audio to text.

Value parameters

input

Audio data to transcribe

options

Configuration for transcription

Attributes

Returns

Result containing Transcription or STTError

Definition Classes

Inherited methods

def isAvailable: Result[Boolean]

Check if this provider is available/healthy. Useful for failover logic and availability checks.

Check if this provider is available/healthy. Useful for failover logic and availability checks.

Attributes

Inherited from:
SpeechToText

Concrete fields

override val name: String

Unique identifier/name of this provider

Unique identifier/name of this provider

Attributes

override val supportedFormats: List[String]

List supported audio formats (e.g., "audio/wav", "audio/mp3")

List supported audio formats (e.g., "audio/wav", "audio/mp3")

Attributes