Upgrading from 0.4.1 to 0.5.0
One page for everything a project has to change to move from 0.4.1 to 0.5.0, grouped by what you use.
0.5.0 is not released yet. The latest release is
v0.4.1. This page describesmainas it will ship, and the artifacts and versions below do not resolve until the release is published. The CHANGELOG’sUnreleasedsection is the complete list; this page is the part of it you have to act on. The per-change notes, with every source break and its reason, are in the Migration Guide.
Table of contents
- In one paragraph
- What do I have to do?
- 1. Dependencies: the change every project makes
- 2. Configuration
- 3. Agent users
- 4. Tracing and metrics
- 5. RAG, memory, MCP, speech and image
- 6. Built-in tools
- 7. Errors, cancellation and types
- 8. Java, Kotlin, Spring, cats-effect and ZIO
- 9. Provider and tracing authors
- What did not change
- Before the release
In one paragraph
In 0.4.1, llm4s-core was one artifact of about 84,000 lines that carried the agent runtime, RAG, MCP, speech,
image generation, the knowledge graph, eleven provider clients and their SDKs. 0.5.0 splits it. llm4s-core is now
the spine (types, errors, configuration, the model and tool APIs, the LLMClient interface and the tracing contract),
and everything else is its own artifact. Most package names stay the same, so for most projects the upgrade is “add the
dependencies you use”, not “change imports”. The extractor exceptions are listed in RAG, memory and the rest;
UsageSummary and ModelUsage move from org.llm4s.agent to org.llm4s.llmconnect.model,
so update explicit imports of those types.
Three other things changed at the same time: the agent runtime was rewritten on a typed graph runtime, API keys are now
read from each vendor’s usual variable, and a number of types were tidied before the compatibility baseline is set.
Security fixes also tightened the built-in tools and the workspace runner, so some calls that worked are now refused.
What do I have to do?
| Your project | Read |
|---|---|
Depends on llm4s-core and calls any model |
Dependencies and Configuration |
Uses Agent, handoffs, guardrails or agent streaming |
Agent users |
Uses PlanRunner, DAG, TypedAgent or CancellationToken |
Orchestration is removed |
| Uses native Ollama with tools | Ollama tool calls |
Uses llm4s-config-policy |
Config-policy patterns |
| Uses Langfuse, OpenTelemetry or Prometheus | Tracing and metrics |
| Uses RAG, memory, MCP, speech, image generation or the knowledge graph | RAG, memory, MCP, speech and image |
Uses the built-in tools (HTTPTool, ShellTool, the file tools) |
Built-in tools |
Uses ContainerisedWorkspace, CodeWorker or the workspace runner image |
Workspace runner commands |
| Matches on error types, retries, or passes durations | Errors, cancellation and types |
| Writes a provider or a tracing backend | Provider and tracing authors |
1. Dependencies: the change every project makes
llm4s-core no longer contains a provider. A project that depended on llm4s-core and called OpenAI now needs
llm4s-openai as well:
1
2
3
4
5
6
7
8
9
// 0.4.1: one artifact carried everything
libraryDependencies += "org.llm4s" %% "llm4s-core" % "0.4.1"
// 0.5.0: the core, plus what you use
libraryDependencies ++= Seq(
"org.llm4s" %% "llm4s-core" % "0.5.0",
"org.llm4s" %% "llm4s-openai" % "0.5.0", // the provider you call
"org.llm4s" %% "llm4s-agent" % "0.5.0" // only if you use Agent
)
For Maven and Gradle the artifact id carries the Scala suffix, as before (llm4s-openai_3):
1
2
3
4
5
<dependency>
<groupId>org.llm4s</groupId>
<artifactId>llm4s-openai_3</artifactId>
<version>0.5.0</version>
</dependency>
1
implementation("org.llm4s:llm4s-openai_3:0.5.0")
In this release every module is published at the same version. Installation has an install line for each one.
Which provider module?
A provider is supplied by the module that carries it. The table is generated from what each module registers (it is
checked by UpgradeGuideSpec), so it is the list of valid provider = "..." values for chat and
EMBEDDING_MODEL=<id>/<model> for embeddings.
| Add | Chat providers (provider = ...) |
Embedding providers |
|---|---|---|
llm4s-openai |
openai, azure, requesty |
openai |
llm4s-anthropic |
anthropic |
|
llm4s-gemini |
gemini, vertexai |
|
llm4s-ollama |
ollama |
ollama |
llm4s-openai-compatible |
openai-compatible, deepseek, zai, openrouter, mistral, cohere |
|
llm4s-bedrock |
bedrock |
|
llm4s-watsonx |
watsonx |
|
llm4s-voyage |
voyage |
|
llm4s-jina |
jina |
|
llm4s-cohere |
cohere |
llm4s-openai-compatible is SDK-free and also covers any endpoint that speaks /chat/completions through the generic
openai-compatible provider, configured entirely from its section. Cohere chat is in that module; llm4s-cohere is
the Cohere embedding provider.
If the module is missing, the failure is at the first use, and it says what to add:
1
2
Provider 'openai' ... is not registered. Registered providers: ... If you expected 'openai',
add the dependency that supplies it, or register it explicitly with ProviderRegistry.of(...).
The other new artifacts
| Add | If you use | Notes |
|---|---|---|
llm4s-agent |
Agent, guardrails, handoffs, the console assistant |
see Agent users |
llm4s-agent-tools |
BuiltinTools and every tool under it |
depends on llm4s-core only, not on llm4s-agent |
llm4s-rag |
RAG, vector stores, chunking, reranking, document extraction | brings llm4s-knowledgegraph with it |
llm4s-knowledgegraph |
the knowledge graph without RAG | |
llm4s-memory |
agent memory | in-memory and SQLite stores |
llm4s-memory-postgres |
PostgresMemoryStore |
brings HikariCP and the Postgres driver |
llm4s-mcp |
the Model Context Protocol client, server and MCPToolRegistry |
|
llm4s-speech |
speech-to-text and text-to-speech | brings Vosk and JNA |
llm4s-image |
image generation and vision | brings llm4s-media |
llm4s-observability |
Langfuse tracing, the trace collector, CostTracker |
|
llm4s-observability-prometheus |
the Prometheus metrics collector and endpoint | |
llm4s-java-api, llm4s-spring-boot-starter, llm4s-effect, llm4s-zio |
new in this release | nothing to migrate if you did not use them |
Unchanged coordinates: llm4s-observability-otel, llm4s-knowledgegraph-neo4j, llm4s-workspace-client and
llm4s-workspace-shared. The relocation POMs at the pre-0.4.0 coordinates (core_3 and the others) are unchanged.
What leaves your classpath
If you relied on llm4s-core to bring any of these, declare them yourself now:
- Apache Tika, POI, PDFBox, jsoup, and the AWS S3 and STS clients (now with
llm4s-rag) - HikariCP and the Postgres JDBC driver (now with
llm4s-memory-postgres) - Vosk and JNA (now with
llm4s-speech) - the Azure OpenAI SDK (replaced by
openai-java, inllm4s-openai) and the Anthropic SDK (inllm4s-anthropic) - the Prometheus client (in
llm4s-observability-prometheus) fansi, which only the console assistant used
llm4s-core now depends on no vendor SDK. The published modules also no longer declare logback-classic and
log4j-to-slf4j, only slf4j-api: if your application relied on the Logback llm4s brought, declare a backend yourself,
or SLF4J logs nothing and prints one “no SLF4J providers were found” warning.
Fat jars
Providers and tracing backends are found with java.util.ServiceLoader. Shading tools overwrite same-named resources by
default, which silently drops every META-INF/services file but one. Configure them to concatenate:
1
2
3
4
5
// sbt-assembly
assembly / assemblyMergeStrategy := {
case PathList("META-INF", "services", _*) => MergeStrategy.filterDistinctLines
case other => (assembly / assemblyMergeStrategy).value(other)
}
1
2
<!-- maven-shade -->
<transformer implementation="org.apache.maven.plugins.shade.resource.ServicesResourceTransformer"/>
If you cannot, register the providers yourself, with no classpath scan at all:
1
2
val registry = ProviderRegistry.ofModules(new Llm4sOpenAIModule).withProvider(BedrockProvider)
LLMConnect.getClient(config)(using registry)
2. Configuration
API keys are bound for you
Each provider module binds its vendor’s usual variable, so with OPENAI_API_KEY exported a section needs only a
provider and a model:
1
2
3
4
5
6
7
8
9
10
llm4s {
providers {
provider = "openai-main"
openai-main {
provider = "openai"
model = "gpt-4o-mini"
}
}
}
This is the snippet DocumentedProviderConfigSpec loads as written. The bound variables:
| Provider | Variable |
|---|---|
openai, requesty, azure |
OPENAI_API_KEY, REQUESTY_API_KEY, AZURE_OPENAI_API_KEY |
anthropic |
ANTHROPIC_API_KEY |
gemini |
GOOGLE_API_KEY, or GEMINI_API_KEY when that is unset |
deepseek, zai, openrouter, mistral, cohere |
DEEPSEEK_API_KEY, ZAI_API_KEY, OPENROUTER_API_KEY, MISTRAL_API_KEY, COHERE_API_KEY |
watsonx, voyage, jina, cohere embeddings |
WATSONX_API_KEY, VOYAGE_API_KEY, JINA_API_KEY, COHERE_API_KEY |
An explicit apiKey = ${?OPENAI_API_KEY} in your section still wins, and a second account’s section
sets its own apiKey. The change in 0.4.1 was that no variable was read unless you bound it yourself; if you did,
your binding keeps working. Azure sections relying on AZURE_API_KEY without an explicit binding
must rename it to AZURE_OPENAI_API_KEY, or retain apiKey = ${?AZURE_API_KEY} in the section.
ollama and bedrock bind no key (Ollama needs none; Bedrock uses AWS credentials).
Smaller configuration changes
- Only the section you load is validated. In
0.4.1every section underllm4s.providerswas validated on every load, so each environment had to fill in every section. Now a section with no key, or whose module is not on the classpath, fails only a load of that section.Llm4sConfig.providers(), which returns all of them, still validates all. - Vertex AI reads
projectandlocation. The oldendpointandorganizationspellings still load, with a warning naming the key to rename them to. - Unknown scalar keys in a section are reported with a warning that lists the keys the provider accepts, which is how a
typo shows up. Scalar extras are ignored, but an undeclared object or list makes configuration loading fail;
remove it or flatten its values into the scalar keys the provider accepts. A key only some providers read
(
organization,endpoint,apiVersion,contextWindow,reserveCompletion) is now declared by those providers; in any other provider’s section it is reported and dropped. reference.confkeys moved with their code. HOCON merges them across jars, so key paths are unchanged. But a build that readsllm4s.rag.*,llm4s.tools.*orllm4s.metrics.*without depending on the module that now carries them no longer gets the defaults.- Provider HTTP timeouts are configurable. A provider or embedding section takes an optional
timeouts { request = 3m, stream = 15m }block (see Timeouts). A section without it behaves as before.timeoutsis now a built-in key of every section, so a provider written outside llm4s can no longer declare an extra of that name. LLM_MODELis not read. That has been true since0.3.2; see the 0.x to 1.0 guide if you are older.
Ollama tool calls
The native Ollama provider now sends supplied tools even when the model registry marks the model as lacking
tool support. A model without that support (such as llama3:latest) returns HTTP 400, mapped to a
ValidationError on tools; the client does not retry without tools. Select a tool-capable model such as
llama3.1, or send no tools.
Config-policy patterns
llm4s-config-policy model and base-URL patterns now match the whole value. A model pattern such as
openai/gpt-4o no longer allows gpt-4o-mini or dated snapshots. To allow dated snapshots, use
openai/gpt-4o(-.*)?. A base-URL pattern that needs to allow paths should end with /.*, for example
https://api\.openai\.com/.*; a bare .* suffix also allows lookalike hosts and should be avoided.
Update patterns that relied on partial matches before upgrading.
The prod / prod-safe preset also requires each key-requiring chat section to set its
own explicit apiKey. Shared llm4s.credentials.<provider>.apiKey is insufficient for
this gate: add apiKey = ${?VAR} to every relevant section, including the default account.
Provider behaviour changes
- Gemini and Vertex AI keep a thinking model’s thought signatures on
AssistantMessage.thinkingand send them back, so a Gemini 3 conversation with tool calls no longer fails with HTTP 400. A thought-summary part is now thinking, not part of the answer; every function call in a streamed chunk is kept (only the first was); and a streamedCompletioncarries its tool calls onCompletion.toolCalls, as OpenAI’s and Ollama’s do. - Z.ai keeps the reasoning you replay: a request that sends an earlier turn’s
reasoning_contentback now also sets"thinking": {"clear_thinking": false}, without which Z.ai’s standard endpoint drops it. Z.ai also honoursCompletionOptions.reasoning, which it used to ignore:ReasoningEffort.Nonenow turns thinking off on the GLM models that allow it, and the other levels are sent in the form the configured model documents. A request without areasoningoption is unchanged; see the CHANGELOG entry for #1681. - OpenAI embeddings with the default base URL (
https://api.openai.com/v1) used to post to/v1/v1/embeddings; they now reach/v1/embeddings. A base URL without/v1(https://api.openai.com, a proxy root) still gets/v1/embeddings, so a workaround of that kind keeps working (#1413). - Streamed tool-call arguments reach
onChunkverbatim. The OpenAI and OpenAI-compatible clients used to parse each argument fragment, so a fragment such as":"lost its quotes. Each now arrives as aujson.Str, as the Anthropic and Bedrock clients’ already did; reassemble them with aStreamingAccumulator. The returnedCompletionwas not affected (#1212). - Anthropic omits
temperaturewhen a thinking budget is set, so extended thinking with the defaultCompletionOptionsno longer fails with HTTP 400. - Voyage now sends
input_type(documentunless the request saysInputPurpose.Query), so re-index for the best retrieval quality; older vectors still work.
3. Agent users
The agent runtime was rewritten on a typed graph runtime. AgentState, AgentEvent and the old loop are gone, and the
agent is built, not constructed with the model call:
1
2
3
4
5
6
7
8
9
10
11
12
13
// 0.4.1
val agent = new Agent(client)
val state = agent.run("query", tools, inputGuardrails = in, outputGuardrails = out, maxSteps = Some(10))
// 0.5.0
val result = for {
agent <- Agent.builder("assistant", client)
.withTools(tools)
.withMiddleware(new GuardrailMiddleware(in, out))
.withMaxSteps(10)
.build()
result <- agent.run("query")
} yield result
UpgradeGuideSpec runs the minimal form of this (Agent.builder(...).build(), then run) against a scripted client.
| 0.4.1 | 0.5.0 |
|---|---|
new Agent(client).run(q, tools, ...) |
Agent.builder(id, client)...build(), then run(q) |
tools, systemMessage, completionOptions, maxSteps, handoffs as arguments |
withTools, and the other builder methods |
inputGuardrails / outputGuardrails arguments |
.withMiddleware(new GuardrailMiddleware(input, output)) |
AgentState |
AgentResult: threadId, runId, activeAgent, status, messages, usage |
state.conversation, state.usageSummary |
result.messages, result.usage |
AgentStatus.Complete |
AgentStatus.Completed(answer); result.answer is the Option[String] |
AgentStatus.Failed(error) |
a Left(GraphError...) from run |
| a guardrail block as an error | AgentStatus.Blocked(guardrail, reason) |
continueConversation(state, ...) |
continueConversation(result, ...), which reads only result.threadId |
runWithEvents, AgentEvent |
agent.stream(threadId, query)(listener), with AgentEvents |
AgentContext (tracing, debug) |
Agent.builder(...).withTracing(tracing) |
Also:
- Threads are kept until forgotten. A thread lives in the agent’s runtime, so call
agent.forget(threadId)when a conversation is over. Handoff(agent)isHandoff.to(id, builder). The target is anAgentBuilder, and its id must equal the handoff id.ToolCallPolicyandPolicyDecisionare removed. A policy becomes anAgentMiddlewareoverridingwrapToolCall;ApprovalMiddlewarecovers approval alone.- Session files:
AgentState.saveToFileandloadFromFileare gone. Saveresult.messages, and pass it back ashistoryon a laterrun. AgentIOandAgentZ(cats-effect and ZIO) wrap the newAgent: configure theAgentBuilderin a function.- Langfuse traces use the run id as the trace id and the thread id as the session id.
- LLM-as-judge guardrails read only a number from 0 to 1.
LLMSafetyGuardrail,LLMFactualityGuardrail,LLMQualityGuardrailandLLMToneGuardrailused to clamp any reply into a score, so a judge answering85or85%passed. Such a reply now fails the guardrail (Could not parse LLM judge score): make the judge answer with only the number. A subclass ofLLMGuardrailthat calls or overridesevaluateWithLLMgets aResult[BigDecimal], not aResult[Double]. A judge guardrail’sthresholdmust now be between 0.0 and 1.0: one outside that range (or NaN), which used to block everything or pass everything, fails everyvalidatewith aValidationErroronthreshold, before the judge is called. Correct the threshold (#1520). - PII masking matches more, and differently.
PIIType.Phonenow also masks international numbers (+44 20 7946 0958) andPIIType.CreditCard15-digit American Express numbers; overlapping matches are merged and replaced once, soPIIMasker.sensitiveandfinancialno longer throw on a 16-digit card number. Digit groups on separate lines are no longer joined into one phone, card or SSN;UTC+5is not a phone number; and the email pattern runs in linear time, where it took seconds to hours on long runs of word characters. The card placeholder is[REDACTED_CARD]. Masked output changes for such inputs; see PII detection and masking and the CHANGELOG entries for #1517 and #1713. - A graph subscription sees every live event.
GraphRuntime.subscribenow joins the live stream before it returns, so progress andStreamEvent.Liveevents sent while it replays the log are no longer lost: they are held up to the subscription’scapacity(beyond it, dropped and reported as aStreamEvent.LiveGap, as for a full queue) and delivered in order among the durable events. No API change (#1731).
Interrupting a blocking call cancels its turn
Agent.run, continueConversation, runMultiTurn, recover and resume return Left(CancelledError) when the
calling thread is interrupted, with the interrupt flag set, and now also cancel the turn they were waiting on. They
return once that turn has ended (within 5 seconds), so recover can follow at once; a caller already interrupted
starts no turn. Before, the turn kept running after run returned. To keep a turn running past an interrupt, use
start, startRecover or startResume and await the AgentRun yourself.
- A turn that had already begun committing its outcome cannot be cancelled. The call then returns that outcome
(
Right,CompletedorSuspended) with the interrupt flag still set. If an interrupt must stop your own code, test the flag, not only the result. run(query)forgets the thread of a failed or cancelled turn. Its random thread id is not carried by aLeft, so once such a turn has ended its thread is forgotten. Name the thread (run(threadId, query)) torecoverit.- Cancelling a graph run also cancels the agent turns its nodes are waiting on.
The Java facade’s blocking JAgent calls go through these, so they change too; see
Java, Kotlin, Spring, cats-effect and ZIO.
Orchestration is removed
org.llm4s.agent.orchestration is deleted: PlanRunner, Plan, Node, Edge, the DAG types, TypedAgent,
Policies, OrchestrationError, MDCContext and CancellationToken, with org.llm4s.types.AgentId and
org.llm4s.types.PlanId from llm4s-core. There is no deprecation period. Build the same flow as a typed graph
(GraphBuilder, run on GraphRuntime), which adds checked state, checkpoints and recovery:
| 0.4.1 | 0.5.0 |
|---|---|
TypedAgent[I, O] and its factories |
a GraphNode[I] given to GraphBuilder.node; call an Agent inside it for an LLM step |
Node, Edge, Plan, Plan.builder |
GraphBuilder.node / edge / staticJoin / dynamicJoin, then compile(entry)(output) |
PlanRunner.execute(plan, inputs, token) |
GraphRuntime.start(threadId, graph, input).flatMap(_.await()) |
PlanRunner(maxConcurrentNodes) |
RunConfig with RunBudgets(maxConcurrency = n) |
Policies.withRetry / withTimeout / withFallback |
retry = RetryPolicy(...) on the node / RunBudgets.withTimeout / ordinary Result code (orElse) |
OrchestrationError |
GraphError |
CancellationToken |
RunHandle.cancel() / AgentRun.cancel(), or interrupting the calling thread |
org.llm4s.types.AgentId / PlanId |
org.llm4s.agent.AgentId / RunId |
The several agents in one graph recipe
(MultiAgentGraphRecipe) is a worked replacement: two specialists in one superstep, then an editor behind a static
join. The Migration Guide
has a before-and-after snippet.
The full mapping, with the event-by-event table, is the Stage 1 note.
4. Tracing and metrics
| 0.4.1 | 0.5.0 |
|---|---|
Langfuse in llm4s-core |
add llm4s-observability; TRACING_MODE=langfuse and the LANGFUSE_* variables are unchanged |
OpenTelemetry in llm4s-observability-otel |
the same artifact; it registers its backend, and the reflection core used to find it is gone |
Prometheus in llm4s-core |
add llm4s-observability-prometheus |
Llm4sConfig.metrics() |
MetricsConfigLoader (public, in llm4s-observability-prometheus) |
Tracing.traceAgentState(state) |
removed; agent runs end with TraceEvent.AgentRunEnded |
A tracing backend is on the classpath exactly when its artifact is, and TRACING_MODE names it. UpgradeGuideSpec
shows llm4s-observability supplying langfuse and the OpenTelemetry backend absent until
llm4s-observability-otel is added. An unknown TRACING_MODE still gives no-op tracing, but now logs an error that
lists the modes available.
5. RAG, memory, MCP, speech and image
Each moved with its package names unchanged. Add the artifact and keep your imports. Source breaks, which are the only places an import or a call changes:
RAG (llm4s-rag)
- The two document extractors became one in
org.llm4s.extract:org.llm4s.rag.extract.DocumentExtractorisorg.llm4s.extract.DocumentExtractor,DefaultDocumentExtractorisTikaDocumentExtractor, andUniversalExtractor.extract(path)isTikaDocumentExtractor.extractFromPath(path), which returns aResult[ExtractedDocument]with the text in.text.ExtractorErrorisProcessingError. EmbeddingClient.encodePath(...)isFileEmbedder.encodeFromPath(path, client, FileEmbeddingConfig(...)).Llm4sConfig.pgSearchIndex()isPgSearchIndexConfigLoader.default().- Deleting or re-syncing a document no longer deletes the chunks of other documents whose id merely starts with the
same text (
doc-1anddoc-10), re-ingesting a document replaces its chunks, and a failed read or listing no longer deletes or clears indexed documents. - Re-chunk what
SentenceChunkerindexed. It used to delete the punctuation, the whitespace and the next sentence’s first letter at every sentence boundary ("Hello world. Next one."became"Hello worldext one."). It now keeps every character, so chunk text and sizes change for any input with a sentence boundary. Indexes built with it (ChunkerFactory.default,"sentence"and the semantic chunker’s fallback included) hold corrupted text: re-chunk and re-embed them (#1718). - The chunkers no longer split a surrogate pair (an emoji, a CJK Extension B character) at a window or overlap
boundary; text without such characters is chunked exactly as before
(#1711).
SimpleChunkerandChunkingUtils.chunkTextno longer throw for a very large window, such astargetSize = Int.MaxValue(#1424).
Embeddings. An EmbeddingRequest says whether its input is a document or a query (InputPurpose), defaulting to
Document. EmbeddingRequest(input, model) compiles and behaves as before; new and .copy do not (see
Errors, cancellation and types). RAG and the vector memory stores now embed their
search text as a query.
Memory (llm4s-memory)
MemoryStore.storeAllis all-or-nothing, and the SQL stores write a batch in one transaction.- Keyword search matches whole words, so it returns fewer memories.
InMemoryStore.search,EmbeddingMemoryStore’s keyword fallback andSimpleMemoryManager.getRelevantContextused to test each query word withString.contains, soimatched “Berlin” andscalamatched “scalability”. A query word must now be a whole word of the memory, split asSQLiteMemoryStore’s FTS5 index splits it (case and Latin accents ignored, punctuation dropped), and there is no stemming:preferno longer matches “Prefers”. If you relied on substring matches, record the word forms you search for, or use an embedding store (#1594). getRelevantContextunder a tightmaxTokensno longer writes a section heading with no memory under it, and stays within the budget; it returns""when no memory fits (#1580).
Image generation. The three overlapping image-format enumerations are now org.llm4s.media.MediaType (in
llm4s-media, which llm4s-image brings).
MCP. A tool result flagged isError is a failure; text results stay text rather than being coerced to JSON; and a
failing tool listing is a Left, not an empty list. Streamable HTTP notifications now send the Accept header the
specification requires, so a server built on an MCP SDK no longer refuses notifications/initialized with
406 Not Acceptable (#1006).
Speech. AudioPreprocessing.resamplePcm16 validates its arguments and returns an exact-length result.
6. Built-in tools
BuiltinTools and every tool under it moved to llm4s-agent-tools, with the same names (coreSafe, withHttpSafe,
withFilesSafe, developmentSafe, customSafe). See the built-in tools guide.
Llm4sConfig.loadBraveSearchTool(),loadDuckDuckGoSearchTool()andloadExaSearchTool()are removed. The same methods are onToolsConfigLoader, inllm4s-agent-tools.ShellConfig.readOnly()no longer allowsenv.- File paths are judged by where they really are. The file tools (
read_file,list_directory,file_info,write_file) resolve every symbolic link and compare the result with each allowed and blocked entry one path component at a time. Allowing/datano longer allows/data-secret(list each directory); a link inside an allowed directory works only if its real target is allowed; a dangling link is refused; andlist_directoryhonoursfollowSymlinks = false. On macOS,/varis a link to/private/var, so the defaultblockedPathsnow also blocks the per-user temporary directoriesjava.io.tmpdirreturns there: set your ownblockedPathsif you read or write there. See Files. - A shell command gets a scrubbed environment. It no longer inherits the process environment, where provider
API keys live: only the variables in the new
ShellConfig.inheritedEnvironment(PATH,LANG,LC_ALL,TERM,SystemRoot) plusenvironment. Name any other variable a command needs (HOME,JAVA_HOME) in one of them;ShellConfig.development()still passes everything. The newpathPolicy, and theShellConfig.readOnlyWithin(policy, workingDirectory)preset, hold a command’s working directory and file-like arguments to aFileConfig;file -C,date -fand similar options that read or write a file the command does not name are refused. See Shell. - More shell commands are refused under a path policy. An option value attached to its flag (
grep -flout,grep --file=lout) is now resolved and checked as a path, like a separate one, so a link out of the allowed directory or a..in it is refused (#1723).sort -o,--output,-T,--temporary-directory,--compress-programand--files0-fromjoin the refused options, in every spelling, for allowlists that addsort; the refused options also match whatever the case of the program name or a.exesuffix (FILE -C). An argument over 4096 characters, or a command whose checks would take more than 20000 file-system lookups, is refused. - On Windows, characters that best-fit to a separator or quote are refused by
ShellTool(and the workspace runner): a fullwidth/or\, a division slash, a yen or won sign, a fullwidth quote, or any character the host’s ANSI code page cannot hold, in an argument, the configuredworkingDirectoryor an environment value. AworkingDirectoryor environment value holding one makes every command fail. POSIX hosts are unchanged (#1766). - The HTTP tool is stricter. The SSRF guard now also refuses IPv6 private and special ranges (
fc00::/7, IPv4-compatible, NAT64, 6to4 and Teredo forms of a blocked IPv4 address, documentation ranges), every IPv6 address outside global unicast2000::/3, and the IPv4 ranges240.0.0.0/4,255.255.255.255,192.0.0.0/24and192.88.99.0/24in every form, withSSRF_BLOCKED; useHttpConfig.withInternalIPsAlloweddeliberately if you need one.HttpConfig.timeoutis now one deadline for the whole call (resolution, connecting, every redirect hop and the body), not a per-read timeout, so size it for the entire download; a zero or negativetimeout, which used to mean “no timeout”, now fails every call at once. The body is read only up tomaxResponseSizebytes. Domain and method names are compared inLocale.ROOT. See HTTP and the CHANGELOG entries for #1408 and #1734. -
A redirect to another origin carries only the headers on
redirectSafeHeaders. WithfollowRedirects, from the first hop that leaves the original origin (scheme, host and port, sohttpstohttpcounts) and on every later hop, the tool sends only the caller-set headers on the newHttpConfig.redirectSafeHeaders(by defaultAccept,Accept-Language,Accept-Encoding,User-Agent, andContent-Typewhen the body is re-sent). It used to strip onlyAuthorization,CookieandProxy-Authorization. A credential header (one of those, a name core’s redaction treats as sensitive, or one ending intokenorkey) is never forwarded, even if listed. A custom header that must follow a cross-origin redirect now has to be listed:1 2 3 4
val http = HttpConfig( followRedirects = true, redirectSafeHeaders = HttpConfig.DefaultRedirectSafeHeaders :+ "X-Request-Id" )
read_fileandwrite_filereturn an error for an unknownencodinginstead of throwing (#1710).get_current_datetimewrites thehumanformat in English on every host (it used the JVM’s locale);timezoneandformatare optional in its schema; and an unsupportedformat, or aformatortimezonethat is not a string, is now an error, where it used to answer in ISO / UTC (#1512).- The calculator returns an error, not
InfinityorNaN, for a result that is not finite (10^400,(-8)^(1/3)). json_toolrefuses a path it cannot read to the end instead of returning the value reached so far: writeitems[0], notitems.[0]. A document nested more than 512 levels deep is refused.
Workspace runner commands
The workspace runner (llm4s-workspace-client, and the runner image) checks an allowlisted command’s arguments, not
only its name, before the process starts (#1715). The new refusals are
plain error codes in the existing response (ARGUMENT_NOT_ALLOWED, PATH_ESCAPE_ATTEMPT,
ENVIRONMENT_NOT_ALLOWED), so the protocol is unchanged, but commands that worked before may now be refused:
- options that write, delete or run other programs (
find -delete/-exec,sort -o,ls -L,grep -R, …); - any argument that is, or resolves to, a location outside the workspace, a
greppattern starting with/included (write[/]apifor/api); gitother than its read subcommands (status,log,show,diff,ls-files,ls-tree,grep,blame,rev-parse, a listingbranch), with most global options refused.gitis confined to the workspace’s own repository: in a workspace that has none but lies inside a larger one it now reportsnot a git repository;environmentvariables other thanLANG,LANGUAGE,LC_*,TZ,TERM,COLUMNS,LINESandNO_COLOR;- on Windows, device names, trailing
.or space, response files,"and the other forms the policy cannot reason about.
Write files through the writeFile / modifyFile operations or the read-write allowlist’s own programs instead. The
full rules are in Command policy (with
On Windows), and the CHANGELOG entry for #1715 lists every refused form.
Commands sent over the WebSocket no longer run through a shell (#1756).
ContainerisedWorkspace.executeCommand and executeCommandWithStreaming used to hand the command string to sh -c
(cmd.exe /c on Windows), so only direct calls on the runner were checked. Both paths now apply every check above,
so a containerised agent’s commands change the most:
- Pipes, redirection,
;,&&,$(...), backquotes and variable expansion are refused (FORBIDDEN_CHARACTERS), not interpreted, and only programs on the runner’s allowlist run. Write each command as one program and its arguments (grep -rn TODO src, notcd src && grep -rn TODO . | head), quote an argument holding spaces, send a second command as a second request, and useworkingDirectoryinstead ofcdandwriteFileinstead of redirection. -
A program the agent needs that is not on the profile’s list, such as a build tool, is added with the runner’s new
WORKSPACE_EXTRA_COMMANDS;ContainerisedWorkspaceandCodeWorkertake it asextraAllowedCommands:1
new ContainerisedWorkspace(workspaceDir, imageName, hostPort, extraAllowedCommands = Set("sbt"))
Shells, interpreters (
python,node,awk,java,make, …), launchers (env,xargs,setsid, …), debuggers, remote and container tools, privilege tools, Windows script hosts and.bat/.cmdnames cannot be added (WorkspaceSandboxConfig.NeverAllowedCommands, matched ignoring case,.exeand a version suffix): naming one, or any name withWORKSPACE_SANDBOX_PROFILE=locked, stops the runner at start-up (#1761). An added program runs any code its arguments allow, so add only one you trust. See Adding programs to the allowlist. - A working directory outside the workspace is now
PATH_ESCAPE_ATTEMPT, notINVALID_DIRECTORY, and a WebSocket command sent with no timeout stops at the sandbox’sdefaultCommandTimeoutinstead of running until it ends. environmentis filtered by the allowlist above on this path too, where it used to be copied into the child.
Programs are started by a trusted absolute path (#1790). The runner
looks an allowlisted program up itself (on Windows in the system and Windows directories, then PATH; on POSIX in
PATH), never in its own or the command’s working directory, never through an empty or relative PATH entry, and
never inside the workspace. A program found only there is refused with EXECUTABLE_NOT_ALLOWED, and one found nowhere
with the new EXECUTABLE_NOT_FOUND (it was EXECUTION_FAILED): match on both if you handled a missing program. A
started program’s PATH keeps only the entries the runner searches, and on Windows cmd.exe built-ins run as
cmd /d /c, so the registry’s AutoRun command no longer runs first.
git runs sandboxed (#1721). Every git the runner starts ignores the
system and global git configuration and attributes (a safe.directory, alias or credential setting in the runner
user’s ~/.gitconfig no longer applies), fetches nothing, and has hooks, core.fsmonitor, external diff, textconv,
filter and merge drivers and signature checking switched off. A repository whose configuration has an
include.path / includeIf, a hook.* key or a driver name outside printable ASCII is refused with
GIT_CONFIG_NOT_ALLOWED; --ignore-submodules, --remerge-diff, --diff-merges and git status -v are refused
with ARGUMENT_NOT_ALLOWED. Writing into .git, in any spelling or through a link, is refused: writeFile /
modifyFile with the new PATH_NOT_ALLOWED, and cp, mv, rm, mkdir, touch, chmod (or an added program)
naming it with ARGUMENT_NOT_ALLOWED. See git’s own files.
Other new refusals:
- An
mvorcpof two sources or more is refused when a later argument goes through a name an earlier source’s operation makes, and a link-preservingcpof two sources whose names the destination file system treats as one (case and Unicode normalisation ignored) is refused. Run one command per source (#1776; see Several sources in one command). sort’s arguments are parsed asgetoptparses them and every file it reads is checked, including every argument after the first operand (which GNUsortmay read as a file), so writesort -t / a.txt, notsort a.txt -t /(#1763; see Option values).- On Windows,
sortandfindstrtake only their native/switches (#1738), and an argument, working directory or environment value holding a character that best-fits to a separator or quote, or one the host’s ANSI code page cannot hold, is refused (ARGUMENT_NOT_ALLOWED,ENVIRONMENT_NOT_ALLOWED; #1766).
Relaxed:
- On POSIX,
rm,unlinkandmvcan remove or rename a symbolic link that points out of the workspace (it wasPATH_ESCAPE_ATTEMPT), when the link is named without a trailing/andrmis not recursive (#1730; see Removing a link). - A command that reads standard input (
catwith no operand,sort,wc,grep xwith no file) gets end-of-file at once instead of hanging until the timeout (#1728).
Removed: WorkspaceConfigSupport.loadSandboxConfig. It read llm4s.workspace.sandbox.profile on the client,
which nothing passed to the runner, so it enforced nothing. What the runner enforces is set only by its own
WORKSPACE_SANDBOX_PROFILE and WORKSPACE_EXTRA_COMMANDS: set them on the runner’s container, and use
WorkspaceSandboxConfig.fromProfileName and WorkspaceSandboxConfig.validate to turn a profile name into a config in
your own code.
The workspace agent protocol lists every error code.
7. Errors, cancellation and types
CancelledErroris a newLLMError, returned when a call is interrupted. Cancellation is thread interruption, orRunHandle.cancel()/AgentRun.cancel()for graph and agent runs;CancellationTokenis removed (see Orchestration is removed).org.llm4s.types.AgentIdandPlanIdare removed fromllm4s-core. The agent’s id isorg.llm4s.agent.AgentId(inllm4s-agent); a run’s isRunId.LLMError.isRecoverable(error)is total and replaces theerror.isRecoverableextension. A custom error is recoverable only if it mixes inRecoverableError; any other is not, and nothing throws aMatchError.- One retry rule. The default retry policy, the client retry wrapper and the graph’s node retry now agree: a recoverable error is retried, except a client-error HTTP status and an optimistic-lock failure.
-
Durations are
FiniteDurationand times areInstant, with the unit gone from the name:timeoutMsistimeout,durationMsisduration. Wire formats and config keys keep their keys and units.1 2
Left(RateLimitError("openai", 1.second)) // was a Long of milliseconds CrawlerConfig(delay = 500.millis, timeout = 30.seconds)
Check string interpolation:
s"${duration}ms"compiles but prints150 millisecondsms; useduration.toMillis. -
Growth-prone types have a private constructor, a public
applywith the defaults, andwith*setters, so they can gain a field without breaking you..copyis no longer available from outside:1
val options = CompletionOptions(temperature = 0.2).withMaxTokens(500) // was .copy(maxTokens = Some(500))
The types are
CompletionOptions,Completion,StreamedChunk,TokenUsage,EmbeddingRequest,ModelCapabilities,ModelMetadata,ProviderConfigSpec,EmbeddingConfigSpec,ProviderFeatures,NamedProviderConfig,ReliabilityConfig,CircuitBreakerConfig,RateLimitConfig,ContextConfig,AssistantMessage,OpenAIConfig,AnthropicConfigandOllamaConfig. Replacemessage.copy(contentOpt = Some(text))withmessage.withContent(text)andconfig.copy(baseUrl = url)withconfig.withBaseUrl(url); usewithToolCalls,withThinking,withApiKeyandwithModelfor the corresponding fields. Replacenewwith companion factories. Java and Kotlin callers useOpenAIConfig.apply(key, model),AnthropicConfig.apply(key, model)orOllamaConfig.apply(model, baseUrl), followed bywith*setters. - Provider config classes gain a trailing
timeoutsfield (OpenAIConfig,AnthropicConfig,OllamaConfig,AzureConfig,GeminiConfigand the others, andEmbeddingProviderConfig). Scala source that callsapplycompiles unchanged; Java and Kotlin callers that pass every argument add the default timeouts last, and a pattern match needs the extra field. See Provider and embedding configs gaintimeouts. Completionhascitations, the sources a model cited (Completion.citations,withCitations). It defaults to empty, so nothing changes for you.- Exceptions mentioning
401or429are no longer guessed to be authentication or rate-limit errors.Safety.safely,Safety.fromTryandtoResultused to turn any exception whose message contained401into anAuthenticationErrorand429into a retryableRateLimitError, dropping the exception. They now classify a status only in an HTTP context (HTTP 401,status code: 429,401 Unauthorized, or a message that starts with401:as the OpenAI and Anthropic SDKs write it); anything else is anUnknownErrorthat keeps the exception as its cause. If you matched on those errors from your own code, check that the message names the status. - JSON a model or a tool produced is refused beyond 512 levels of nesting. A reply, tool-call arguments or a
structured-output document nested that deep is a
Left(or a malformed tool call), not aStackOverflowError. - Caches.
CacheKeyGenerator.sha256now length-prefixes its parts, and the key function ofCachedEmbeddingClienttakes(text, modelName, purpose: InputPurpose); the default isCacheKeyGenerator.embeddingKey. A custom key function gains the parameter and should include it in the key. Vectors stored under the old keys in a persistentEmbeddingCacheare not found and are re-embedded on first use.CacheConfig.createrefuses aNaNsimilarity threshold. - Logged request and response bodies are redacted in more shapes (credentials under compound keys, in JSON inside
a string, in arrays and objects, in
key=valuepairs), and a provider’s error body is now redacted before it is cut short for a log line or an error message, so a key straddling the cut no longer leaks as a fragment.Proxy-Authorization,Cookie,Set-CookieandX-Amz-Security-Tokenare now sensitive in every shape, and a cookie header’s whole value is replaced;cookieis matched as a whole word, socookie_policyis kept (#1686). Values in backslash-escaped quotes, JSON-escaped&,=and?(\u0026,\u003d,\u003f), andkey=valuelogs serialised into JSON strings are read too. Logs and error messages that showed a session cookie or such a value now show[REDACTED]; nothing else to change. See provider exchange logging for what is and is not redacted. Llm4sHttpClientreturnsResultand takesFiniteDuration.- Removed, not deprecated. Before the compatibility baseline, dead, deprecated and accidental public API was
deleted: for example
Result.fromTry(uset.toResult),ToolBuilder#build()(usebuildSafe()),ToolRegistry#getToolDefinitionsSafe(usegetOpenAITools()),ReliableProvidersandReliabilitySyntax. The passes are listed, with their replacements, in the Migration Guide.
8. Java, Kotlin, Spring, cats-effect and ZIO
llm4s-java-api, llm4s-spring-boot-starter, llm4s-effect and llm4s-zio are new artifacts, and a Kotlin coroutine
API is in the repository (Experimental). None of them was in 0.4.1, so there is nothing to migrate from it. If you use
AgentIO or AgentZ, see Agent users.
If you built against a snapshot of main, these changed before the release:
- Agent results are Java types.
JAgent.run,continueConversation,resume,recover,AgentStream.await()andAgentStreamListener.onCompletereturn aJAgentResult, not the ScalaAgentResult:answer()is anOptional<String>,messages()ajava.util.List<JMessage>,usage()aJUsageSummary, andstatus().kind()the Java enumAgentStatusKind. The KotlinAgentKtfunctions andAgentStreamItem.Donetake or carry aJAgentResulttoo. Replaceimport org.llm4s.agent.AgentResultwithorg.llm4s.javaapi.JAgentResult; the CHANGELOG entry for #1393 has the full mapping. - Suspended turns.
JAgent.pending(result),resume(threadId, answers)andrecover(threadId)(and the KotlinAgentKt.pending,resumeandrecover) handle a turn waiting for approval; see Suspended turns from Java and Kotlin (#1392).JAgent.pending(result)is now a shortcut forresult.status().pending(). - Kotlin cancellation. Cancelling
AgentKt.runorcontinueConversation(a cancelled scope,withTimeout) now cancels the turn, asresume,recoverand the streams do, and returns once it has ended, leaving the thread forrecover; it used to stop only the wait, while the turn kept running and the thread stayed busy. - Interrupts. The Java facade never throws
InterruptedException: an interrupted call is aCancelledErrorresult, with the thread’s interrupt flag set. Test the result rather than catching the exception. - An interrupted
JAgentcall cancels its turn. The blockingJAgent.run,continueConversation,resumeandrecoverused to stop only the wait and leave the turn running; they now cancel it and return once it has ended, leaving the thread forrecover, as the Scala calls and Kotlin do (see Interrupting a blocking call cancels its turn). A turn that had already begun committing returns its real outcome with the interrupt flag set.AgentStream.cancel()still cancels a streamed turn without interrupting any thread, and now waits up to 5 seconds for it to end, as the blocking calls do, instead of without a bound. - Forgetting threads. Kotlin
AgentKt.run(query), like ScalaAgent.run(query)and JavaJAgent.run(query), forgets the random thread of a turn that failed or was cancelled; it used to stay in the agent’s runtime. The newJAgent.forget(threadId)forgets a conversation by its thread id, for one whose turn left no result (#1682, #1688). - Whole replies, options and typed errors.
JLlmClient.completion(...)returns anLlmResult<JCompletion>with the content, model, tool calls, token usage and estimated cost;completestill returns the text.JCompletionOptions.builder()sets temperature, top-p, token limits, penalties andJReasoningEffortwithoutscala.Option.LlmException.getKind()gives anLlmErrorKind, withisRecoverable(),getRetryAfter()andgetStatusCode(): callisRecoverable()rather than infer it from the kind (#1487, #1488). - Embeddings.
Llm4s.createDefaultEmbeddingClient()returns aJEmbeddingClientfor the modelEMBEDDING_MODELnames.LlmException.getKind()now reads an embedding provider’sEmbeddingErrorby its HTTP status (AUTHENTICATION,RATE_LIMIT,VALIDATION,SERVICE) where it was alwaysOTHER(#1490). See the Java guide. - These modules need JDK 21, like the rest of llm4s.
9. Provider and tracing authors
- A provider is a
ProviderDescriptorin its own module, supplied by anLlm4sProviderModuledeclared inMETA-INF/services/org.llm4s.llmconnect.spi.Llm4sProviderModule. It must be aclasswith a public no-argument constructor, not anobject:ServiceLoaderinstantiates the named class. - Adding a provider is adding a dependency.
ProviderRegistry.defaultisProviderRegistry.discover(), computed once. A broken services entry is recorded as a failure and does not take out the others. ProviderKindbecame the opaque typeProviderId, andProviderConfigis open to providers outside core. ThefromValuesfactories (OpenAIConfig.fromValues,AzureConfig.fromValuesand the others) return aResultinstead of throwing.- Provider-specific keys are declared in
ProviderConfigSpec.extrasand read fromNamedProviderConfig.extras. - A tracing backend implements
org.llm4s.trace.spi.TracingBackend, declared inMETA-INF/services/org.llm4s.trace.spi.TracingBackend, withmode = TracingMode.Named("yourmode"). llm4s-provider-testkit(ProviderModuleChecks) is how a provider module proves its registration.
See writing a provider.
What did not change
- Most package names, apart from the extractor consolidation and the
UsageSummary/ModelUsagemove above. - Scala 3 (3.7.1) and JDK 21.
- Named provider sections under
llm4s.providers, andLlm4sConfig.defaultProvider(). - The tool API:
ToolFunction,ToolRegistry,ToolBuilderandSchemastay inllm4s-core, so your own tools need nothing new. Agent.run(query)as the way to run an agent.
Before the release
For the maintainer who cuts 0.5.0. This page was written against main on 2026-10-07, brought up to date on
2026-10-10 and again on 2026-10-11 (through #1806 and #1722), and has to be
re-checked:
- Version. The snippets show
0.5.0literally, on purpose (migration guides are exempt fromscripts/check-doc-versions.sh). Change them if the number changes. - Artifacts. Run
sbt -error listPublishedArtifactsandscripts/verify-release.sh, and compare the output with the tables in Dependencies. - Anything merged since. Read the CHANGELOG’s
Unreleasedsection for entries added after this page, and add what a user has to act on. - The tests.
sbt "samples/testOnly org.llm4s.samples.upgrade.UpgradeGuideSpec"checks the provider table, explicit registration, the agent snippet, an interruptedrun, the removed orchestration, tracing discovery, the renamed RAG and error APIs, theredirectSafeHeaderssnippet and whole-word memory search.sbt "core/testOnly org.llm4s.llmconnect.spi.CoreOnlyProviderDiscoverySpec"checks the core-only provider boundary. - Links. Run
scripts/check-doc-links.shandscripts/check-doc-support.sh. - The install guide. Remove the “Not yet published” notes in
docs/getting-started/installation.md, and update the g8 starter template, which still pins0.4.0.