An org.llm4s.llmconnect.LLMClient that answers from a script, for testing code that calls LLM4S without a model, a network or an API key.
Two ways to script it:
- ScriptedLLMClient.sequence answers the first call with the first reply, the second call with the second, and so on. This suits an agent loop: a tool call, then the final answer.
- ScriptedLLMClient.respondingTo chooses the reply from the conversation, so the answer does not depend on how many calls came before.
val client = ScriptedLLMClient.sequence(
Reply.toolCall("get_weather", """{"city": "Paris"}"""),
Reply.text("It is sunny in Paris.")
)
// ... run the code under test with `client` ...
client.callCount // 2
client.calls.head.lastUserText
Every request is recorded (calls), whether it came through complete or streamComplete. A call the script has no reply for does not hang or throw: it returns a Left(UnscriptedCallError) that names the call and shows the last message, so a test that makes more calls than it scripted fails with a message that says why.
streamComplete replays the same reply as chunks: the text in pieces of at most chunkSize characters, then each tool call, then a final chunk that carries the finish reason (stop, or tool_calls when the reply calls tools).
The client is thread-safe: calls from several threads are recorded one at a time and each gets its own position in the script. It reports a context window of getContextWindow tokens and reserves getReserveCompletion for the answer, which withContextWindow and withReserveCompletion change.
Attributes
- Companion
- object
- Graph
-
- Supertypes
Members list
Value members
Concrete methods
The number of requests received so far.
The number of requests received so far.
Attributes
Every request received so far, oldest first.
Every request received so far, oldest first.
Attributes
Executes a blocking completion request and returns the full response.
Executes a blocking completion request and returns the full response.
Sends the conversation to the LLM and waits for the complete response. Use when you need the entire response at once or when streaming is not required.
Value parameters
- conversation
-
conversation history including system, user, assistant, and tool messages
- options
-
configuration including temperature, max tokens, tools, etc. (default: CompletionOptions())
Attributes
- Returns
-
Right(Completion) with the model's response, or Left(LLMError) on failure
- Definition Classes
Returns the maximum context window size supported by this model in tokens.
Returns the maximum context window size supported by this model in tokens.
The context window is the total tokens (prompt + completion) the model can process in a single request, including all conversation messages and the generated response.
Attributes
- Returns
-
total context window size in tokens (e.g., 4096, 8192, 128000)
- Definition Classes
Returns the number of tokens reserved for the model's completion response.
Returns the number of tokens reserved for the model's completion response.
This value is subtracted from the context window when calculating available tokens for prompts. Corresponds to the max_tokens or completion token limit configured for the model.
Attributes
- Returns
-
number of tokens reserved for completion
- Definition Classes
The most recent request, if there was one.
The most recent request, if there was one.
Attributes
Executes a streaming completion request, invoking a callback for each chunk as it arrives.
Executes a streaming completion request, invoking a callback for each chunk as it arrives.
Streams the response incrementally, calling onChunk for each token/chunk received. Enables real-time display of responses. Returns the final accumulated completion on success.
Value parameters
- conversation
-
conversation history including system, user, assistant, and tool messages
- onChunk
-
callback invoked for each chunk; called synchronously, avoid blocking operations
- options
-
configuration including temperature, max tokens, tools, etc. (default: CompletionOptions())
Attributes
- Returns
-
Right(Completion) with the complete accumulated response, or Left(LLMError) on failure
- Definition Classes
A fresh client with the same script and no recorded calls, streaming text in chunks of this size (at least 1).
A fresh client with the same script and no recorded calls, streaming text in chunks of this size (at least 1).
Attributes
A fresh client with the same script and no recorded calls, reporting another context window.
A fresh client with the same script and no recorded calls, reporting another context window.
Attributes
A fresh client with the same script and no recorded calls, reserving another number of tokens for the answer.
A fresh client with the same script and no recorded calls, reserving another number of tokens for the answer.
Attributes
Inherited methods
Releases resources and closes connections to the LLM provider.
Releases resources and closes connections to the LLM provider.
Call when the client is no longer needed. After calling close(), the client should not be used. Default implementation is a no-op; override if managing resources like connections or thread pools.
Attributes
- Inherited from:
- LLMClient
Sends the conversation and parses the response into a typed value using the provided schema.
Sends the conversation and parses the response into a typed value using the provided schema.
Sets ResponseFormat.JsonSchema on the options. OpenAI, Azure OpenAI and Gemini enforce the schema at generation time, as does Ollama 0.5 or later through its format field. Requesty and the OpenAI-compatible providers (including Cohere) send it as response_format, but whether it is enforced is up to the server - for Requesty, a router, the backend model it routes to: one that ignores the field returns unconstrained text. Anthropic falls back to a best-effort system-prompt instruction, which is not schema-enforced. Clients that do not read responseFormat (watsonx, Bedrock) send no schema at all. Because models may wrap JSON in markdown code fences or surround it with prose, the response is normalised (fence stripped, first balanced {...} or [...] extracted) before being deserialised with uPickle into the expected type A.
The reply is '''not''' validated against the schema: it is only deserialised. A constraint the reader does not check - an enum, a numeric or string bound, additionalProperties = false (extra keys are ignored) - can be violated by a reply that still returns Right(A). Check such constraints on the result yourself.
The schema is derived with strict = true, which lists '''every''' property as required, including a property declared optional with required = false. Only responseFormat is overridden: every other option you pass is forwarded unchanged to complete, where the provider client may adjust or drop options the model does not support, as for any other complete call. name and strict on the format are left at their defaults ("response" and true); call complete with your own ResponseFormat.JsonSchema to set them.
Type parameters
- A
-
target type; must have a corresponding
upickle.default.Reader[A]
Value parameters
- conversation
-
conversation history
- options
-
additional completion options (default: CompletionOptions())
- reader
-
implicit uPickle reader for deserialising the JSON into
A - schema
-
JSON-Schema description of the expected response object
Attributes
- Returns
-
Right(A) on success. Left(ValidationError) with field
structured_outputwhen the reply is not JSON, is JSONnull, or cannot be deserialised asA; any other Left is the provider call's own error, returned unchanged - Inherited from:
- LLMClient
Calculates available token budget for prompts after accounting for completion reserve and headroom.
Calculates available token budget for prompts after accounting for completion reserve and headroom.
Formula: (contextWindow - reserveCompletion) * (1 - headroom)
Headroom provides a safety margin for tokenization variations and message formatting overhead.
Value parameters
- headroom
-
safety margin as percentage of prompt budget (default: HeadroomPercent.Standard ~10%)
Attributes
- Returns
-
maximum tokens available for prompt content
- Inherited from:
- LLMClient
Validates client configuration and connectivity to the LLM provider.
Validates client configuration and connectivity to the LLM provider.
May perform checks such as verifying API credentials, testing connectivity, and validating configuration. Default implementation returns success; override for provider-specific validation.
Attributes
- Returns
-
Right(()) if validation succeeds, Left(LLMError) with details on failure
- Inherited from:
- LLMClient