LLMClient for IBM watsonx.ai text generation (POST /ml/v1/text/generation and /ml/v1/text/generation_stream).
'''Beta - built on deprecated endpoints.''' IBM's February 2026 release notes (https://www.ibm.com/docs/en/software-hub/5.3.x?topic=new-watsonxai) deprecate the watsonx.ai "Infer text" and "Infer text event stream" endpoints (/ml/v1/text/generation and /generation_stream) this module uses; IBM points to the chat API. This module has never been run against the live service (no watsonx account), its API is not frozen, and tools are unsupported because of the endpoint. Migration to the chat API: https://github.com/llm4s/llm4s/issues/1314.
== Authentication ==
The IBM Cloud API key is exchanged at the IAM endpoint for a bearer token that lives an hour. The token is cached and refreshed lazily, five minutes before it expires.
== Request format ==
The text-generation API is not a chat API: the conversation is flattened into one input string with [SYSTEM]:, [USER]:, [ASSISTANT]: and [TOOL_RESULT:<id>]: prefixes, ending in an open [ASSISTANT]: turn. Content is not escaped, so user content can forge those markers (a prompt-injection surface inherent to the flattened format). Requests carry stop_sequences (WatsonXClient.StopSequences) so a model cannot go on to write the next turn itself.
== Unsupported options ==
'''Tools are rejected.''' text-generation has no tool calling: complete and streamComplete return a Left(ValidationError("tools", ...)) when CompletionOptions.tools is non-empty, before any HTTP call (the IAM exchange included).
'''Ignored without error:''' presencePenalty, frequencyPenalty, responseFormat, reasoning and budgetTokens. Only temperature, maxTokens and topP (when not 1.0) are sent.
== Stream endings ==
A stream must end with a terminal event (a stop_reason other than not_finished). One that ends without it, or whose reason is in WatsonXClient.ErrorStopReasons, is a Left(ServiceError) naming the reason; text received so far is not returned as a success. Every other reason (eos_token, stop_sequence, max_tokens, token_limit, unknown values) is a normal stop. complete applies the same rule to results[0].stop_reason (a missing one is fine).
Value parameters
config
model, credentials, project or space and endpoints.
exchangeLogging
optional provider exchange logging.
httpClient
used for both the model calls and the IAM exchange.
Executes a blocking completion request and returns the full response.
Executes a blocking completion request and returns the full response.
Sends the conversation to the LLM and waits for the complete response. Use when you need the entire response at once or when streaming is not required.
Value parameters
conversation
conversation history including system, user, assistant, and tool messages
options
configuration including temperature, max tokens, tools, etc. (default: CompletionOptions())
Attributes
Returns
Right(Completion) with the model's response, or Left(LLMError) on failure
Returns the maximum context window size supported by this model in tokens.
Returns the maximum context window size supported by this model in tokens.
The context window is the total tokens (prompt + completion) the model can process in a single request, including all conversation messages and the generated response.
Attributes
Returns
total context window size in tokens (e.g., 4096, 8192, 128000)
Returns the number of tokens reserved for the model's completion response.
Returns the number of tokens reserved for the model's completion response.
This value is subtracted from the context window when calculating available tokens for prompts. Corresponds to the max_tokens or completion token limit configured for the model.
Executes a streaming completion request, invoking a callback for each chunk as it arrives.
Executes a streaming completion request, invoking a callback for each chunk as it arrives.
Streams the response incrementally, calling onChunk for each token/chunk received. Enables real-time display of responses. Returns the final accumulated completion on success.
Value parameters
conversation
conversation history including system, user, assistant, and tool messages
onChunk
callback invoked for each chunk; called synchronously, avoid blocking operations
options
configuration including temperature, max tokens, tools, etc. (default: CompletionOptions())
Attributes
Returns
Right(Completion) with the complete accumulated response, or Left(LLMError) on failure
Releases resources and closes connections to the LLM provider.
Releases resources and closes connections to the LLM provider.
Call when the client is no longer needed. After calling close(), the client should not be used. Default implementation is a no-op; override if managing resources like connections or thread pools.
Sends the conversation and parses the response into a typed value using the provided schema.
Sends the conversation and parses the response into a typed value using the provided schema.
Sets ResponseFormat.JsonSchema on the options so providers that support native structured output (OpenAI, Gemini) enforce the schema at generation time. Anthropic falls back to a best-effort system-prompt instruction, which is not schema-enforced. Because models may wrap JSON in markdown code fences or surround it with prose, the response is normalised (fence stripped, first balanced {...} or [...] extracted) before being deserialised with uPickle into the expected type A.
Type parameters
A
target type; must have a corresponding upickle.default.Reader[A]
Validates that the client is open, executes the operation, and records standard completion metrics (latency, token usage, estimated cost).
Validates that the client is open, executes the operation, and records standard completion metrics (latency, token usage, estimated cost).
Use this in complete and streamComplete implementations to avoid repeating the lifecycle-check + metrics-wrapping boilerplate.
An interrupted call - one that throws InterruptedException, or fails while the thread is interrupted - is returned as Left(CancelledError) with the interrupt flag kept, whatever the provider SDK did with it. A call made with the flag already set returns Left(CancelledError) at once, without running operation: an SDK that ignores the flag would otherwise send the (billed) request anyway.
Value parameters
operation
The provider-specific completion logic to execute. Called only when the client is open and the thread is not interrupted.
Attributes
Returns
The completion result with metrics recorded as a side-effect.
Validates client configuration and connectivity to the LLM provider.
Validates client configuration and connectivity to the LLM provider.
May perform checks such as verifying API credentials, testing connectivity, and validating configuration. Default implementation returns success; override for provider-specific validation.
Attributes
Returns
Right(()) if validation succeeds, Left(LLMError) with details on failure
protected def withMetrics[A](provider: String, model: String, operation: => Result[A], extractUsage: A => Option[TokenUsage], extractCost: A => Option[Double]): Result[A]
Executes operation and records metrics for the call.
Executes operation and records metrics for the call.
Latency and outcome (success or classified error) are recorded for every call regardless of result. Token counts and cost are recorded only on success — a Left result emits an org.llm4s.metrics.Outcome.Error event whose kind is derived from the org.llm4s.error.LLMError subtype via ErrorKind.fromLLMError.
Value parameters
extractCost
Extracts the pre-computed cost (USD) from a successful result; return None to skip cost recording.
extractUsage
Extracts prompt/completion token counts from a successful result; return None to skip token recording.
model
Model identifier forwarded to the collector.
operation
The LLM call to time and observe.
provider
Provider label forwarded to the collector (e.g. "openai").