ReliableClient

org.llm4s.reliability.ReliableClient
@Stable
final class ReliableClient(underlying: LLMClient, providerName: String, config: ReliabilityConfig, collector: Option[MetricsCollector], clock: () => Instant, sleep: FiniteDuration => Unit) extends LLMClient

Wrapper that adds reliability features to any LLMClient.

Provides:

  • Retry with configurable policies (exponential backoff, linear, fixed)
  • Circuit breaker to fail fast when service is down
  • Deadline enforcement to prevent hanging operations
  • Local token-bucket rate limiting (RateLimitConfig), checked before every attempt
  • Metrics tracking for retry attempts and circuit breaker state

Thread-safety: Uses AtomicInteger/AtomicReference for circuit breaker state management to ensure correct behavior under concurrent access.

Value parameters

clock

the current time, read for deadlines and circuit-breaker recovery; injectable for tests

collector

Optional metrics collector for observability

config

Reliability configuration

providerName

Explicit provider name for stable metrics labels

sleep

how the client waits between attempts; injectable so a test can record the delays chosen, or throw InterruptedException to simulate an interrupted wait

underlying

The client to wrap

Attributes

Graph
Supertypes
trait LLMClient
trait AutoCloseable
class Object
trait Matchable
class Any

Members list

Value members

Concrete methods

override def close(): Unit

Releases resources and closes connections to the LLM provider.

Releases resources and closes connections to the LLM provider.

Call when the client is no longer needed. After calling close(), the client should not be used. Default implementation is a no-op; override if managing resources like connections or thread pools.

Attributes

Definition Classes
LLMClient -> AutoCloseable
override def complete(conversation: Conversation, options: CompletionOptions): Result[Completion]

Executes a blocking completion request and returns the full response.

Executes a blocking completion request and returns the full response.

Sends the conversation to the LLM and waits for the complete response. Use when you need the entire response at once or when streaming is not required.

Value parameters

conversation

conversation history including system, user, assistant, and tool messages

options

configuration including temperature, max tokens, tools, etc. (default: CompletionOptions())

Attributes

Returns

Right(Completion) with the model's response, or Left(LLMError) on failure

Definition Classes

Get current circuit breaker state (for testing/monitoring).

Get current circuit breaker state (for testing/monitoring).

Attributes

override def getContextWindow(): Int

Returns the maximum context window size supported by this model in tokens.

Returns the maximum context window size supported by this model in tokens.

The context window is the total tokens (prompt + completion) the model can process in a single request, including all conversation messages and the generated response.

Attributes

Returns

total context window size in tokens (e.g., 4096, 8192, 128000)

Definition Classes
override def getReserveCompletion(): Int

Returns the number of tokens reserved for the model's completion response.

Returns the number of tokens reserved for the model's completion response.

This value is subtracted from the context window when calculating available tokens for prompts. Corresponds to the max_tokens or completion token limit configured for the model.

Attributes

Returns

number of tokens reserved for completion

Definition Classes
def resetCircuitBreaker(): Unit

Reset circuit breaker state (for testing).

Reset circuit breaker state (for testing).

Attributes

override def streamComplete(conversation: Conversation, options: CompletionOptions, onChunk: StreamedChunk => Unit): Result[Completion]

Executes a streaming completion request, invoking a callback for each chunk as it arrives.

Executes a streaming completion request, invoking a callback for each chunk as it arrives.

Streams the response incrementally, calling onChunk for each token/chunk received. Enables real-time display of responses. Returns the final accumulated completion on success.

Value parameters

conversation

conversation history including system, user, assistant, and tool messages

onChunk

callback invoked for each chunk; called synchronously, avoid blocking operations

options

configuration including temperature, max tokens, tools, etc. (default: CompletionOptions())

Attributes

Returns

Right(Completion) with the complete accumulated response, or Left(LLMError) on failure

Definition Classes
override def validate(): Result[Unit]

Validates client configuration and connectivity to the LLM provider.

Validates client configuration and connectivity to the LLM provider.

May perform checks such as verifying API credentials, testing connectivity, and validating configuration. Default implementation returns success; override for provider-specific validation.

Attributes

Returns

Right(()) if validation succeeds, Left(LLMError) with details on failure

Definition Classes

Inherited methods

def completeStructured[A](conversation: Conversation, schema: ObjectSchema[A], options: CompletionOptions)(implicit reader: Reader[A]): Result[A]

Sends the conversation and parses the response into a typed value using the provided schema.

Sends the conversation and parses the response into a typed value using the provided schema.

Sets ResponseFormat.JsonSchema on the options. OpenAI, Azure OpenAI and Gemini enforce the schema at generation time, as does Ollama 0.5 or later through its format field. Requesty and the OpenAI-compatible providers (including Cohere) send it as response_format, but whether it is enforced is up to the server - for Requesty, a router, the backend model it routes to: one that ignores the field returns unconstrained text. Anthropic falls back to a best-effort system-prompt instruction, which is not schema-enforced. Clients that do not read responseFormat (watsonx, Bedrock) send no schema at all. Because models may wrap JSON in markdown code fences or surround it with prose, the response is normalised (fence stripped, first balanced {...} or [...] extracted) before being deserialised with uPickle into the expected type A.

The reply is '''not''' validated against the schema: it is only deserialised. A constraint the reader does not check - an enum, a numeric or string bound, additionalProperties = false (extra keys are ignored) - can be violated by a reply that still returns Right(A). Check such constraints on the result yourself.

The schema is derived with strict = true, which lists '''every''' property as required, including a property declared optional with required = false. Only responseFormat is overridden: every other option you pass is forwarded unchanged to complete, where the provider client may adjust or drop options the model does not support, as for any other complete call. name and strict on the format are left at their defaults ("response" and true); call complete with your own ResponseFormat.JsonSchema to set them.

Type parameters

A

target type; must have a corresponding upickle.default.Reader[A]

Value parameters

conversation

conversation history

options

additional completion options (default: CompletionOptions())

reader

implicit uPickle reader for deserialising the JSON into A

schema

JSON-Schema description of the expected response object

Attributes

Returns

Right(A) on success. Left(ValidationError) with field structured_output when the reply is not JSON, is JSON null, or cannot be deserialised as A; any other Left is the provider call's own error, returned unchanged

Inherited from:
LLMClient

Calculates available token budget for prompts after accounting for completion reserve and headroom.

Calculates available token budget for prompts after accounting for completion reserve and headroom.

Formula: (contextWindow - reserveCompletion) * (1 - headroom)

Headroom provides a safety margin for tokenization variations and message formatting overhead.

Value parameters

headroom

safety margin as percentage of prompt budget (default: HeadroomPercent.Standard ~10%)

Attributes

Returns

maximum tokens available for prompt content

Inherited from:
LLMClient