Upgrading from 0.4.1 to 0.5.0

One page for everything a project has to change to move from 0.4.1 to 0.5.0, grouped by what you use.

0.5.0 is not released yet. The latest release is v0.4.1. This page describes main as it will ship, and the artifacts and versions below do not resolve until the release is published. The CHANGELOG’s Unreleased section is the complete list; this page is the part of it you have to act on. The per-change notes, with every source break and its reason, are in the Migration Guide.

Table of contents

  1. In one paragraph
  2. What do I have to do?
  3. 1. Dependencies: the change every project makes
    1. Which provider module?
    2. The other new artifacts
    3. What leaves your classpath
    4. Fat jars
  4. 2. Configuration
    1. API keys are bound for you
    2. Smaller configuration changes
    3. Ollama tool calls
    4. Config-policy patterns
    5. Provider behaviour changes
  5. 3. Agent users
    1. Interrupting a blocking call cancels its turn
    2. Orchestration is removed
  6. 4. Tracing and metrics
  7. 5. RAG, memory, MCP, speech and image
  8. 6. Built-in tools
    1. Workspace runner commands
  9. 7. Errors, cancellation and types
  10. 8. Java, Kotlin, Spring, cats-effect and ZIO
  11. 9. Provider and tracing authors
  12. What did not change
  13. Before the release

In one paragraph

In 0.4.1, llm4s-core was one artifact of about 84,000 lines that carried the agent runtime, RAG, MCP, speech, image generation, the knowledge graph, eleven provider clients and their SDKs. 0.5.0 splits it. llm4s-core is now the spine (types, errors, configuration, the model and tool APIs, the LLMClient interface and the tracing contract), and everything else is its own artifact. Most package names stay the same, so for most projects the upgrade is “add the dependencies you use”, not “change imports”. The extractor exceptions are listed in RAG, memory and the rest; UsageSummary and ModelUsage move from org.llm4s.agent to org.llm4s.llmconnect.model, so update explicit imports of those types. Three other things changed at the same time: the agent runtime was rewritten on a typed graph runtime, API keys are now read from each vendor’s usual variable, and a number of types were tidied before the compatibility baseline is set. Security fixes also tightened the built-in tools and the workspace runner, so some calls that worked are now refused.

What do I have to do?

Your project Read
Depends on llm4s-core and calls any model Dependencies and Configuration
Uses Agent, handoffs, guardrails or agent streaming Agent users
Uses PlanRunner, DAG, TypedAgent or CancellationToken Orchestration is removed
Uses native Ollama with tools Ollama tool calls
Uses llm4s-config-policy Config-policy patterns
Uses Langfuse, OpenTelemetry or Prometheus Tracing and metrics
Uses RAG, memory, MCP, speech, image generation or the knowledge graph RAG, memory, MCP, speech and image
Uses the built-in tools (HTTPTool, ShellTool, the file tools) Built-in tools
Uses ContainerisedWorkspace, CodeWorker or the workspace runner image Workspace runner commands
Matches on error types, retries, or passes durations Errors, cancellation and types
Writes a provider or a tracing backend Provider and tracing authors

1. Dependencies: the change every project makes

llm4s-core no longer contains a provider. A project that depended on llm4s-core and called OpenAI now needs llm4s-openai as well:

1
2
3
4
5
6
7
8
9
// 0.4.1: one artifact carried everything
libraryDependencies += "org.llm4s" %% "llm4s-core" % "0.4.1"

// 0.5.0: the core, plus what you use
libraryDependencies ++= Seq(
  "org.llm4s" %% "llm4s-core"   % "0.5.0",
  "org.llm4s" %% "llm4s-openai" % "0.5.0", // the provider you call
  "org.llm4s" %% "llm4s-agent"  % "0.5.0"  // only if you use Agent
)

For Maven and Gradle the artifact id carries the Scala suffix, as before (llm4s-openai_3):

1
2
3
4
5
<dependency>
  <groupId>org.llm4s</groupId>
  <artifactId>llm4s-openai_3</artifactId>
  <version>0.5.0</version>
</dependency>
1
implementation("org.llm4s:llm4s-openai_3:0.5.0")

In this release every module is published at the same version. Installation has an install line for each one.

Which provider module?

A provider is supplied by the module that carries it. The table is generated from what each module registers (it is checked by UpgradeGuideSpec), so it is the list of valid provider = "..." values for chat and EMBEDDING_MODEL=<id>/<model> for embeddings.

Add Chat providers (provider = ...) Embedding providers
llm4s-openai openai, azure, requesty openai
llm4s-anthropic anthropic  
llm4s-gemini gemini, vertexai  
llm4s-ollama ollama ollama
llm4s-openai-compatible openai-compatible, deepseek, zai, openrouter, mistral, cohere  
llm4s-bedrock bedrock  
llm4s-watsonx watsonx  
llm4s-voyage   voyage
llm4s-jina   jina
llm4s-cohere   cohere

llm4s-openai-compatible is SDK-free and also covers any endpoint that speaks /chat/completions through the generic openai-compatible provider, configured entirely from its section. Cohere chat is in that module; llm4s-cohere is the Cohere embedding provider.

If the module is missing, the failure is at the first use, and it says what to add:

1
2
Provider 'openai' ... is not registered. Registered providers: ... If you expected 'openai',
add the dependency that supplies it, or register it explicitly with ProviderRegistry.of(...).

The other new artifacts

Add If you use Notes
llm4s-agent Agent, guardrails, handoffs, the console assistant see Agent users
llm4s-agent-tools BuiltinTools and every tool under it depends on llm4s-core only, not on llm4s-agent
llm4s-rag RAG, vector stores, chunking, reranking, document extraction brings llm4s-knowledgegraph with it
llm4s-knowledgegraph the knowledge graph without RAG  
llm4s-memory agent memory in-memory and SQLite stores
llm4s-memory-postgres PostgresMemoryStore brings HikariCP and the Postgres driver
llm4s-mcp the Model Context Protocol client, server and MCPToolRegistry  
llm4s-speech speech-to-text and text-to-speech brings Vosk and JNA
llm4s-image image generation and vision brings llm4s-media
llm4s-observability Langfuse tracing, the trace collector, CostTracker  
llm4s-observability-prometheus the Prometheus metrics collector and endpoint  
llm4s-java-api, llm4s-spring-boot-starter, llm4s-effect, llm4s-zio new in this release nothing to migrate if you did not use them

Unchanged coordinates: llm4s-observability-otel, llm4s-knowledgegraph-neo4j, llm4s-workspace-client and llm4s-workspace-shared. The relocation POMs at the pre-0.4.0 coordinates (core_3 and the others) are unchanged.

What leaves your classpath

If you relied on llm4s-core to bring any of these, declare them yourself now:

  • Apache Tika, POI, PDFBox, jsoup, and the AWS S3 and STS clients (now with llm4s-rag)
  • HikariCP and the Postgres JDBC driver (now with llm4s-memory-postgres)
  • Vosk and JNA (now with llm4s-speech)
  • the Azure OpenAI SDK (replaced by openai-java, in llm4s-openai) and the Anthropic SDK (in llm4s-anthropic)
  • the Prometheus client (in llm4s-observability-prometheus)
  • fansi, which only the console assistant used

llm4s-core now depends on no vendor SDK. The published modules also no longer declare logback-classic and log4j-to-slf4j, only slf4j-api: if your application relied on the Logback llm4s brought, declare a backend yourself, or SLF4J logs nothing and prints one “no SLF4J providers were found” warning.

Fat jars

Providers and tracing backends are found with java.util.ServiceLoader. Shading tools overwrite same-named resources by default, which silently drops every META-INF/services file but one. Configure them to concatenate:

1
2
3
4
5
// sbt-assembly
assembly / assemblyMergeStrategy := {
  case PathList("META-INF", "services", _*) => MergeStrategy.filterDistinctLines
  case other                                => (assembly / assemblyMergeStrategy).value(other)
}
1
2
<!-- maven-shade -->
<transformer implementation="org.apache.maven.plugins.shade.resource.ServicesResourceTransformer"/>

If you cannot, register the providers yourself, with no classpath scan at all:

1
2
val registry = ProviderRegistry.ofModules(new Llm4sOpenAIModule).withProvider(BedrockProvider)
LLMConnect.getClient(config)(using registry)

2. Configuration

API keys are bound for you

Each provider module binds its vendor’s usual variable, so with OPENAI_API_KEY exported a section needs only a provider and a model:

1
2
3
4
5
6
7
8
9
10
llm4s {
  providers {
    provider = "openai-main"

    openai-main {
      provider = "openai"
      model    = "gpt-4o-mini"
    }
  }
}

This is the snippet DocumentedProviderConfigSpec loads as written. The bound variables:

Provider Variable
openai, requesty, azure OPENAI_API_KEY, REQUESTY_API_KEY, AZURE_OPENAI_API_KEY
anthropic ANTHROPIC_API_KEY
gemini GOOGLE_API_KEY, or GEMINI_API_KEY when that is unset
deepseek, zai, openrouter, mistral, cohere DEEPSEEK_API_KEY, ZAI_API_KEY, OPENROUTER_API_KEY, MISTRAL_API_KEY, COHERE_API_KEY
watsonx, voyage, jina, cohere embeddings WATSONX_API_KEY, VOYAGE_API_KEY, JINA_API_KEY, COHERE_API_KEY

An explicit apiKey = ${?OPENAI_API_KEY} in your section still wins, and a second account’s section sets its own apiKey. The change in 0.4.1 was that no variable was read unless you bound it yourself; if you did, your binding keeps working. Azure sections relying on AZURE_API_KEY without an explicit binding must rename it to AZURE_OPENAI_API_KEY, or retain apiKey = ${?AZURE_API_KEY} in the section. ollama and bedrock bind no key (Ollama needs none; Bedrock uses AWS credentials).

Smaller configuration changes

  • Only the section you load is validated. In 0.4.1 every section under llm4s.providers was validated on every load, so each environment had to fill in every section. Now a section with no key, or whose module is not on the classpath, fails only a load of that section. Llm4sConfig.providers(), which returns all of them, still validates all.
  • Vertex AI reads project and location. The old endpoint and organization spellings still load, with a warning naming the key to rename them to.
  • Unknown scalar keys in a section are reported with a warning that lists the keys the provider accepts, which is how a typo shows up. Scalar extras are ignored, but an undeclared object or list makes configuration loading fail; remove it or flatten its values into the scalar keys the provider accepts. A key only some providers read (organization, endpoint, apiVersion, contextWindow, reserveCompletion) is now declared by those providers; in any other provider’s section it is reported and dropped.
  • reference.conf keys moved with their code. HOCON merges them across jars, so key paths are unchanged. But a build that reads llm4s.rag.*, llm4s.tools.* or llm4s.metrics.* without depending on the module that now carries them no longer gets the defaults.
  • Provider HTTP timeouts are configurable. A provider or embedding section takes an optional timeouts { request = 3m, stream = 15m } block (see Timeouts). A section without it behaves as before. timeouts is now a built-in key of every section, so a provider written outside llm4s can no longer declare an extra of that name.
  • LLM_MODEL is not read. That has been true since 0.3.2; see the 0.x to 1.0 guide if you are older.

Ollama tool calls

The native Ollama provider now sends supplied tools even when the model registry marks the model as lacking tool support. A model without that support (such as llama3:latest) returns HTTP 400, mapped to a ValidationError on tools; the client does not retry without tools. Select a tool-capable model such as llama3.1, or send no tools.

Config-policy patterns

llm4s-config-policy model and base-URL patterns now match the whole value. A model pattern such as openai/gpt-4o no longer allows gpt-4o-mini or dated snapshots. To allow dated snapshots, use openai/gpt-4o(-.*)?. A base-URL pattern that needs to allow paths should end with /.*, for example https://api\.openai\.com/.*; a bare .* suffix also allows lookalike hosts and should be avoided. Update patterns that relied on partial matches before upgrading.

The prod / prod-safe preset also requires each key-requiring chat section to set its own explicit apiKey. Shared llm4s.credentials.<provider>.apiKey is insufficient for this gate: add apiKey = ${?VAR} to every relevant section, including the default account.

Provider behaviour changes

  • Gemini and Vertex AI keep a thinking model’s thought signatures on AssistantMessage.thinking and send them back, so a Gemini 3 conversation with tool calls no longer fails with HTTP 400. A thought-summary part is now thinking, not part of the answer; every function call in a streamed chunk is kept (only the first was); and a streamed Completion carries its tool calls on Completion.toolCalls, as OpenAI’s and Ollama’s do.
  • Z.ai keeps the reasoning you replay: a request that sends an earlier turn’s reasoning_content back now also sets "thinking": {"clear_thinking": false}, without which Z.ai’s standard endpoint drops it. Z.ai also honours CompletionOptions.reasoning, which it used to ignore: ReasoningEffort.None now turns thinking off on the GLM models that allow it, and the other levels are sent in the form the configured model documents. A request without a reasoning option is unchanged; see the CHANGELOG entry for #1681.
  • OpenAI embeddings with the default base URL (https://api.openai.com/v1) used to post to /v1/v1/embeddings; they now reach /v1/embeddings. A base URL without /v1 (https://api.openai.com, a proxy root) still gets /v1/embeddings, so a workaround of that kind keeps working (#1413).
  • Streamed tool-call arguments reach onChunk verbatim. The OpenAI and OpenAI-compatible clients used to parse each argument fragment, so a fragment such as ":" lost its quotes. Each now arrives as a ujson.Str, as the Anthropic and Bedrock clients’ already did; reassemble them with a StreamingAccumulator. The returned Completion was not affected (#1212).
  • Anthropic omits temperature when a thinking budget is set, so extended thinking with the default CompletionOptions no longer fails with HTTP 400.
  • Voyage now sends input_type (document unless the request says InputPurpose.Query), so re-index for the best retrieval quality; older vectors still work.

3. Agent users

The agent runtime was rewritten on a typed graph runtime. AgentState, AgentEvent and the old loop are gone, and the agent is built, not constructed with the model call:

1
2
3
4
5
6
7
8
9
10
11
12
13
// 0.4.1
val agent = new Agent(client)
val state = agent.run("query", tools, inputGuardrails = in, outputGuardrails = out, maxSteps = Some(10))

// 0.5.0
val result = for {
  agent  <- Agent.builder("assistant", client)
              .withTools(tools)
              .withMiddleware(new GuardrailMiddleware(in, out))
              .withMaxSteps(10)
              .build()
  result <- agent.run("query")
} yield result

UpgradeGuideSpec runs the minimal form of this (Agent.builder(...).build(), then run) against a scripted client.

0.4.1 0.5.0
new Agent(client).run(q, tools, ...) Agent.builder(id, client)...build(), then run(q)
tools, systemMessage, completionOptions, maxSteps, handoffs as arguments withTools, and the other builder methods
inputGuardrails / outputGuardrails arguments .withMiddleware(new GuardrailMiddleware(input, output))
AgentState AgentResult: threadId, runId, activeAgent, status, messages, usage
state.conversation, state.usageSummary result.messages, result.usage
AgentStatus.Complete AgentStatus.Completed(answer); result.answer is the Option[String]
AgentStatus.Failed(error) a Left(GraphError...) from run
a guardrail block as an error AgentStatus.Blocked(guardrail, reason)
continueConversation(state, ...) continueConversation(result, ...), which reads only result.threadId
runWithEvents, AgentEvent agent.stream(threadId, query)(listener), with AgentEvents
AgentContext (tracing, debug) Agent.builder(...).withTracing(tracing)

Also:

  • Threads are kept until forgotten. A thread lives in the agent’s runtime, so call agent.forget(threadId) when a conversation is over.
  • Handoff(agent) is Handoff.to(id, builder). The target is an AgentBuilder, and its id must equal the handoff id.
  • ToolCallPolicy and PolicyDecision are removed. A policy becomes an AgentMiddleware overriding wrapToolCall; ApprovalMiddleware covers approval alone.
  • Session files: AgentState.saveToFile and loadFromFile are gone. Save result.messages, and pass it back as history on a later run.
  • AgentIO and AgentZ (cats-effect and ZIO) wrap the new Agent: configure the AgentBuilder in a function.
  • Langfuse traces use the run id as the trace id and the thread id as the session id.
  • LLM-as-judge guardrails read only a number from 0 to 1. LLMSafetyGuardrail, LLMFactualityGuardrail, LLMQualityGuardrail and LLMToneGuardrail used to clamp any reply into a score, so a judge answering 85 or 85% passed. Such a reply now fails the guardrail (Could not parse LLM judge score): make the judge answer with only the number. A subclass of LLMGuardrail that calls or overrides evaluateWithLLM gets a Result[BigDecimal], not a Result[Double]. A judge guardrail’s threshold must now be between 0.0 and 1.0: one outside that range (or NaN), which used to block everything or pass everything, fails every validate with a ValidationError on threshold, before the judge is called. Correct the threshold (#1520).
  • PII masking matches more, and differently. PIIType.Phone now also masks international numbers (+44 20 7946 0958) and PIIType.CreditCard 15-digit American Express numbers; overlapping matches are merged and replaced once, so PIIMasker.sensitive and financial no longer throw on a 16-digit card number. Digit groups on separate lines are no longer joined into one phone, card or SSN; UTC+5 is not a phone number; and the email pattern runs in linear time, where it took seconds to hours on long runs of word characters. The card placeholder is [REDACTED_CARD]. Masked output changes for such inputs; see PII detection and masking and the CHANGELOG entries for #1517 and #1713.
  • A graph subscription sees every live event. GraphRuntime.subscribe now joins the live stream before it returns, so progress and StreamEvent.Live events sent while it replays the log are no longer lost: they are held up to the subscription’s capacity (beyond it, dropped and reported as a StreamEvent.LiveGap, as for a full queue) and delivered in order among the durable events. No API change (#1731).

Interrupting a blocking call cancels its turn

Agent.run, continueConversation, runMultiTurn, recover and resume return Left(CancelledError) when the calling thread is interrupted, with the interrupt flag set, and now also cancel the turn they were waiting on. They return once that turn has ended (within 5 seconds), so recover can follow at once; a caller already interrupted starts no turn. Before, the turn kept running after run returned. To keep a turn running past an interrupt, use start, startRecover or startResume and await the AgentRun yourself.

  • A turn that had already begun committing its outcome cannot be cancelled. The call then returns that outcome (Right, Completed or Suspended) with the interrupt flag still set. If an interrupt must stop your own code, test the flag, not only the result.
  • run(query) forgets the thread of a failed or cancelled turn. Its random thread id is not carried by a Left, so once such a turn has ended its thread is forgotten. Name the thread (run(threadId, query)) to recover it.
  • Cancelling a graph run also cancels the agent turns its nodes are waiting on.

The Java facade’s blocking JAgent calls go through these, so they change too; see Java, Kotlin, Spring, cats-effect and ZIO.

Orchestration is removed

org.llm4s.agent.orchestration is deleted: PlanRunner, Plan, Node, Edge, the DAG types, TypedAgent, Policies, OrchestrationError, MDCContext and CancellationToken, with org.llm4s.types.AgentId and org.llm4s.types.PlanId from llm4s-core. There is no deprecation period. Build the same flow as a typed graph (GraphBuilder, run on GraphRuntime), which adds checked state, checkpoints and recovery:

0.4.1 0.5.0
TypedAgent[I, O] and its factories a GraphNode[I] given to GraphBuilder.node; call an Agent inside it for an LLM step
Node, Edge, Plan, Plan.builder GraphBuilder.node / edge / staticJoin / dynamicJoin, then compile(entry)(output)
PlanRunner.execute(plan, inputs, token) GraphRuntime.start(threadId, graph, input).flatMap(_.await())
PlanRunner(maxConcurrentNodes) RunConfig with RunBudgets(maxConcurrency = n)
Policies.withRetry / withTimeout / withFallback retry = RetryPolicy(...) on the node / RunBudgets.withTimeout / ordinary Result code (orElse)
OrchestrationError GraphError
CancellationToken RunHandle.cancel() / AgentRun.cancel(), or interrupting the calling thread
org.llm4s.types.AgentId / PlanId org.llm4s.agent.AgentId / RunId

The several agents in one graph recipe (MultiAgentGraphRecipe) is a worked replacement: two specialists in one superstep, then an editor behind a static join. The Migration Guide has a before-and-after snippet.

The full mapping, with the event-by-event table, is the Stage 1 note.


4. Tracing and metrics

0.4.1 0.5.0
Langfuse in llm4s-core add llm4s-observability; TRACING_MODE=langfuse and the LANGFUSE_* variables are unchanged
OpenTelemetry in llm4s-observability-otel the same artifact; it registers its backend, and the reflection core used to find it is gone
Prometheus in llm4s-core add llm4s-observability-prometheus
Llm4sConfig.metrics() MetricsConfigLoader (public, in llm4s-observability-prometheus)
Tracing.traceAgentState(state) removed; agent runs end with TraceEvent.AgentRunEnded

A tracing backend is on the classpath exactly when its artifact is, and TRACING_MODE names it. UpgradeGuideSpec shows llm4s-observability supplying langfuse and the OpenTelemetry backend absent until llm4s-observability-otel is added. An unknown TRACING_MODE still gives no-op tracing, but now logs an error that lists the modes available.


5. RAG, memory, MCP, speech and image

Each moved with its package names unchanged. Add the artifact and keep your imports. Source breaks, which are the only places an import or a call changes:

RAG (llm4s-rag)

  • The two document extractors became one in org.llm4s.extract: org.llm4s.rag.extract.DocumentExtractor is org.llm4s.extract.DocumentExtractor, DefaultDocumentExtractor is TikaDocumentExtractor, and UniversalExtractor.extract(path) is TikaDocumentExtractor.extractFromPath(path), which returns a Result[ExtractedDocument] with the text in .text. ExtractorError is ProcessingError.
  • EmbeddingClient.encodePath(...) is FileEmbedder.encodeFromPath(path, client, FileEmbeddingConfig(...)).
  • Llm4sConfig.pgSearchIndex() is PgSearchIndexConfigLoader.default().
  • Deleting or re-syncing a document no longer deletes the chunks of other documents whose id merely starts with the same text (doc-1 and doc-10), re-ingesting a document replaces its chunks, and a failed read or listing no longer deletes or clears indexed documents.
  • Re-chunk what SentenceChunker indexed. It used to delete the punctuation, the whitespace and the next sentence’s first letter at every sentence boundary ("Hello world. Next one." became "Hello worldext one."). It now keeps every character, so chunk text and sizes change for any input with a sentence boundary. Indexes built with it (ChunkerFactory.default, "sentence" and the semantic chunker’s fallback included) hold corrupted text: re-chunk and re-embed them (#1718).
  • The chunkers no longer split a surrogate pair (an emoji, a CJK Extension B character) at a window or overlap boundary; text without such characters is chunked exactly as before (#1711). SimpleChunker and ChunkingUtils.chunkText no longer throw for a very large window, such as targetSize = Int.MaxValue (#1424).

Embeddings. An EmbeddingRequest says whether its input is a document or a query (InputPurpose), defaulting to Document. EmbeddingRequest(input, model) compiles and behaves as before; new and .copy do not (see Errors, cancellation and types). RAG and the vector memory stores now embed their search text as a query.

Memory (llm4s-memory)

  • MemoryStore.storeAll is all-or-nothing, and the SQL stores write a batch in one transaction.
  • Keyword search matches whole words, so it returns fewer memories. InMemoryStore.search, EmbeddingMemoryStore’s keyword fallback and SimpleMemoryManager.getRelevantContext used to test each query word with String.contains, so i matched “Berlin” and scala matched “scalability”. A query word must now be a whole word of the memory, split as SQLiteMemoryStore’s FTS5 index splits it (case and Latin accents ignored, punctuation dropped), and there is no stemming: prefer no longer matches “Prefers”. If you relied on substring matches, record the word forms you search for, or use an embedding store (#1594).
  • getRelevantContext under a tight maxTokens no longer writes a section heading with no memory under it, and stays within the budget; it returns "" when no memory fits (#1580).

Image generation. The three overlapping image-format enumerations are now org.llm4s.media.MediaType (in llm4s-media, which llm4s-image brings).

MCP. A tool result flagged isError is a failure; text results stay text rather than being coerced to JSON; and a failing tool listing is a Left, not an empty list. Streamable HTTP notifications now send the Accept header the specification requires, so a server built on an MCP SDK no longer refuses notifications/initialized with 406 Not Acceptable (#1006).

Speech. AudioPreprocessing.resamplePcm16 validates its arguments and returns an exact-length result.


6. Built-in tools

BuiltinTools and every tool under it moved to llm4s-agent-tools, with the same names (coreSafe, withHttpSafe, withFilesSafe, developmentSafe, customSafe). See the built-in tools guide.

  • Llm4sConfig.loadBraveSearchTool(), loadDuckDuckGoSearchTool() and loadExaSearchTool() are removed. The same methods are on ToolsConfigLoader, in llm4s-agent-tools.
  • ShellConfig.readOnly() no longer allows env.
  • File paths are judged by where they really are. The file tools (read_file, list_directory, file_info, write_file) resolve every symbolic link and compare the result with each allowed and blocked entry one path component at a time. Allowing /data no longer allows /data-secret (list each directory); a link inside an allowed directory works only if its real target is allowed; a dangling link is refused; and list_directory honours followSymlinks = false. On macOS, /var is a link to /private/var, so the default blockedPaths now also blocks the per-user temporary directories java.io.tmpdir returns there: set your own blockedPaths if you read or write there. See Files.
  • A shell command gets a scrubbed environment. It no longer inherits the process environment, where provider API keys live: only the variables in the new ShellConfig.inheritedEnvironment (PATH, LANG, LC_ALL, TERM, SystemRoot) plus environment. Name any other variable a command needs (HOME, JAVA_HOME) in one of them; ShellConfig.development() still passes everything. The new pathPolicy, and the ShellConfig.readOnlyWithin(policy, workingDirectory) preset, hold a command’s working directory and file-like arguments to a FileConfig; file -C, date -f and similar options that read or write a file the command does not name are refused. See Shell.
  • More shell commands are refused under a path policy. An option value attached to its flag (grep -flout, grep --file=lout) is now resolved and checked as a path, like a separate one, so a link out of the allowed directory or a .. in it is refused (#1723). sort -o, --output, -T, --temporary-directory, --compress-program and --files0-from join the refused options, in every spelling, for allowlists that add sort; the refused options also match whatever the case of the program name or a .exe suffix (FILE -C). An argument over 4096 characters, or a command whose checks would take more than 20000 file-system lookups, is refused.
  • On Windows, characters that best-fit to a separator or quote are refused by ShellTool (and the workspace runner): a fullwidth / or \, a division slash, a yen or won sign, a fullwidth quote, or any character the host’s ANSI code page cannot hold, in an argument, the configured workingDirectory or an environment value. A workingDirectory or environment value holding one makes every command fail. POSIX hosts are unchanged (#1766).
  • The HTTP tool is stricter. The SSRF guard now also refuses IPv6 private and special ranges (fc00::/7, IPv4-compatible, NAT64, 6to4 and Teredo forms of a blocked IPv4 address, documentation ranges), every IPv6 address outside global unicast 2000::/3, and the IPv4 ranges 240.0.0.0/4, 255.255.255.255, 192.0.0.0/24 and 192.88.99.0/24 in every form, with SSRF_BLOCKED; use HttpConfig.withInternalIPsAllowed deliberately if you need one. HttpConfig.timeout is now one deadline for the whole call (resolution, connecting, every redirect hop and the body), not a per-read timeout, so size it for the entire download; a zero or negative timeout, which used to mean “no timeout”, now fails every call at once. The body is read only up to maxResponseSize bytes. Domain and method names are compared in Locale.ROOT. See HTTP and the CHANGELOG entries for #1408 and #1734.
  • A redirect to another origin carries only the headers on redirectSafeHeaders. With followRedirects, from the first hop that leaves the original origin (scheme, host and port, so https to http counts) and on every later hop, the tool sends only the caller-set headers on the new HttpConfig.redirectSafeHeaders (by default Accept, Accept-Language, Accept-Encoding, User-Agent, and Content-Type when the body is re-sent). It used to strip only Authorization, Cookie and Proxy-Authorization. A credential header (one of those, a name core’s redaction treats as sensitive, or one ending in token or key) is never forwarded, even if listed. A custom header that must follow a cross-origin redirect now has to be listed:

    1
    2
    3
    4
    
    val http = HttpConfig(
      followRedirects = true,
      redirectSafeHeaders = HttpConfig.DefaultRedirectSafeHeaders :+ "X-Request-Id"
    )
    
  • read_file and write_file return an error for an unknown encoding instead of throwing (#1710).
  • get_current_datetime writes the human format in English on every host (it used the JVM’s locale); timezone and format are optional in its schema; and an unsupported format, or a format or timezone that is not a string, is now an error, where it used to answer in ISO / UTC (#1512).
  • The calculator returns an error, not Infinity or NaN, for a result that is not finite (10^400, (-8)^(1/3)).
  • json_tool refuses a path it cannot read to the end instead of returning the value reached so far: write items[0], not items.[0]. A document nested more than 512 levels deep is refused.

Workspace runner commands

The workspace runner (llm4s-workspace-client, and the runner image) checks an allowlisted command’s arguments, not only its name, before the process starts (#1715). The new refusals are plain error codes in the existing response (ARGUMENT_NOT_ALLOWED, PATH_ESCAPE_ATTEMPT, ENVIRONMENT_NOT_ALLOWED), so the protocol is unchanged, but commands that worked before may now be refused:

  • options that write, delete or run other programs (find -delete / -exec, sort -o, ls -L, grep -R, …);
  • any argument that is, or resolves to, a location outside the workspace, a grep pattern starting with / included (write [/]api for /api);
  • git other than its read subcommands (status, log, show, diff, ls-files, ls-tree, grep, blame, rev-parse, a listing branch), with most global options refused. git is confined to the workspace’s own repository: in a workspace that has none but lies inside a larger one it now reports not a git repository;
  • environment variables other than LANG, LANGUAGE, LC_*, TZ, TERM, COLUMNS, LINES and NO_COLOR;
  • on Windows, device names, trailing . or space, response files, " and the other forms the policy cannot reason about.

Write files through the writeFile / modifyFile operations or the read-write allowlist’s own programs instead. The full rules are in Command policy (with On Windows), and the CHANGELOG entry for #1715 lists every refused form.

Commands sent over the WebSocket no longer run through a shell (#1756). ContainerisedWorkspace.executeCommand and executeCommandWithStreaming used to hand the command string to sh -c (cmd.exe /c on Windows), so only direct calls on the runner were checked. Both paths now apply every check above, so a containerised agent’s commands change the most:

  • Pipes, redirection, ;, &&, $(...), backquotes and variable expansion are refused (FORBIDDEN_CHARACTERS), not interpreted, and only programs on the runner’s allowlist run. Write each command as one program and its arguments (grep -rn TODO src, not cd src && grep -rn TODO . | head), quote an argument holding spaces, send a second command as a second request, and use workingDirectory instead of cd and writeFile instead of redirection.
  • A program the agent needs that is not on the profile’s list, such as a build tool, is added with the runner’s new WORKSPACE_EXTRA_COMMANDS; ContainerisedWorkspace and CodeWorker take it as extraAllowedCommands:

    1
    
    new ContainerisedWorkspace(workspaceDir, imageName, hostPort, extraAllowedCommands = Set("sbt"))
    

    Shells, interpreters (python, node, awk, java, make, …), launchers (env, xargs, setsid, …), debuggers, remote and container tools, privilege tools, Windows script hosts and .bat / .cmd names cannot be added (WorkspaceSandboxConfig.NeverAllowedCommands, matched ignoring case, .exe and a version suffix): naming one, or any name with WORKSPACE_SANDBOX_PROFILE=locked, stops the runner at start-up (#1761). An added program runs any code its arguments allow, so add only one you trust. See Adding programs to the allowlist.

  • A working directory outside the workspace is now PATH_ESCAPE_ATTEMPT, not INVALID_DIRECTORY, and a WebSocket command sent with no timeout stops at the sandbox’s defaultCommandTimeout instead of running until it ends.
  • environment is filtered by the allowlist above on this path too, where it used to be copied into the child.

Programs are started by a trusted absolute path (#1790). The runner looks an allowlisted program up itself (on Windows in the system and Windows directories, then PATH; on POSIX in PATH), never in its own or the command’s working directory, never through an empty or relative PATH entry, and never inside the workspace. A program found only there is refused with EXECUTABLE_NOT_ALLOWED, and one found nowhere with the new EXECUTABLE_NOT_FOUND (it was EXECUTION_FAILED): match on both if you handled a missing program. A started program’s PATH keeps only the entries the runner searches, and on Windows cmd.exe built-ins run as cmd /d /c, so the registry’s AutoRun command no longer runs first.

git runs sandboxed (#1721). Every git the runner starts ignores the system and global git configuration and attributes (a safe.directory, alias or credential setting in the runner user’s ~/.gitconfig no longer applies), fetches nothing, and has hooks, core.fsmonitor, external diff, textconv, filter and merge drivers and signature checking switched off. A repository whose configuration has an include.path / includeIf, a hook.* key or a driver name outside printable ASCII is refused with GIT_CONFIG_NOT_ALLOWED; --ignore-submodules, --remerge-diff, --diff-merges and git status -v are refused with ARGUMENT_NOT_ALLOWED. Writing into .git, in any spelling or through a link, is refused: writeFile / modifyFile with the new PATH_NOT_ALLOWED, and cp, mv, rm, mkdir, touch, chmod (or an added program) naming it with ARGUMENT_NOT_ALLOWED. See git’s own files.

Other new refusals:

  • An mv or cp of two sources or more is refused when a later argument goes through a name an earlier source’s operation makes, and a link-preserving cp of two sources whose names the destination file system treats as one (case and Unicode normalisation ignored) is refused. Run one command per source (#1776; see Several sources in one command).
  • sort’s arguments are parsed as getopt parses them and every file it reads is checked, including every argument after the first operand (which GNU sort may read as a file), so write sort -t / a.txt, not sort a.txt -t / (#1763; see Option values).
  • On Windows, sort and findstr take only their native / switches (#1738), and an argument, working directory or environment value holding a character that best-fits to a separator or quote, or one the host’s ANSI code page cannot hold, is refused (ARGUMENT_NOT_ALLOWED, ENVIRONMENT_NOT_ALLOWED; #1766).

Relaxed:

  • On POSIX, rm, unlink and mv can remove or rename a symbolic link that points out of the workspace (it was PATH_ESCAPE_ATTEMPT), when the link is named without a trailing / and rm is not recursive (#1730; see Removing a link).
  • A command that reads standard input (cat with no operand, sort, wc, grep x with no file) gets end-of-file at once instead of hanging until the timeout (#1728).

Removed: WorkspaceConfigSupport.loadSandboxConfig. It read llm4s.workspace.sandbox.profile on the client, which nothing passed to the runner, so it enforced nothing. What the runner enforces is set only by its own WORKSPACE_SANDBOX_PROFILE and WORKSPACE_EXTRA_COMMANDS: set them on the runner’s container, and use WorkspaceSandboxConfig.fromProfileName and WorkspaceSandboxConfig.validate to turn a profile name into a config in your own code. The workspace agent protocol lists every error code.


7. Errors, cancellation and types

  • CancelledError is a new LLMError, returned when a call is interrupted. Cancellation is thread interruption, or RunHandle.cancel() / AgentRun.cancel() for graph and agent runs; CancellationToken is removed (see Orchestration is removed).
  • org.llm4s.types.AgentId and PlanId are removed from llm4s-core. The agent’s id is org.llm4s.agent.AgentId (in llm4s-agent); a run’s is RunId.
  • LLMError.isRecoverable(error) is total and replaces the error.isRecoverable extension. A custom error is recoverable only if it mixes in RecoverableError; any other is not, and nothing throws a MatchError.
  • One retry rule. The default retry policy, the client retry wrapper and the graph’s node retry now agree: a recoverable error is retried, except a client-error HTTP status and an optimistic-lock failure.
  • Durations are FiniteDuration and times are Instant, with the unit gone from the name: timeoutMs is timeout, durationMs is duration. Wire formats and config keys keep their keys and units.

    1
    2
    
    Left(RateLimitError("openai", 1.second))      // was a Long of milliseconds
    CrawlerConfig(delay = 500.millis, timeout = 30.seconds)
    

    Check string interpolation: s"${duration}ms" compiles but prints 150 millisecondsms; use duration.toMillis.

  • Growth-prone types have a private constructor, a public apply with the defaults, and with* setters, so they can gain a field without breaking you. .copy is no longer available from outside:

    1
    
    val options = CompletionOptions(temperature = 0.2).withMaxTokens(500) // was .copy(maxTokens = Some(500))
    

    The types are CompletionOptions, Completion, StreamedChunk, TokenUsage, EmbeddingRequest, ModelCapabilities, ModelMetadata, ProviderConfigSpec, EmbeddingConfigSpec, ProviderFeatures, NamedProviderConfig, ReliabilityConfig, CircuitBreakerConfig, RateLimitConfig, ContextConfig, AssistantMessage, OpenAIConfig, AnthropicConfig and OllamaConfig. Replace message.copy(contentOpt = Some(text)) with message.withContent(text) and config.copy(baseUrl = url) with config.withBaseUrl(url); use withToolCalls, withThinking, withApiKey and withModel for the corresponding fields. Replace new with companion factories. Java and Kotlin callers use OpenAIConfig.apply(key, model), AnthropicConfig.apply(key, model) or OllamaConfig.apply(model, baseUrl), followed by with* setters.

  • Provider config classes gain a trailing timeouts field (OpenAIConfig, AnthropicConfig, OllamaConfig, AzureConfig, GeminiConfig and the others, and EmbeddingProviderConfig). Scala source that calls apply compiles unchanged; Java and Kotlin callers that pass every argument add the default timeouts last, and a pattern match needs the extra field. See Provider and embedding configs gain timeouts.
  • Completion has citations, the sources a model cited (Completion.citations, withCitations). It defaults to empty, so nothing changes for you.
  • Exceptions mentioning 401 or 429 are no longer guessed to be authentication or rate-limit errors. Safety.safely, Safety.fromTry and toResult used to turn any exception whose message contained 401 into an AuthenticationError and 429 into a retryable RateLimitError, dropping the exception. They now classify a status only in an HTTP context (HTTP 401, status code: 429, 401 Unauthorized, or a message that starts with 401: as the OpenAI and Anthropic SDKs write it); anything else is an UnknownError that keeps the exception as its cause. If you matched on those errors from your own code, check that the message names the status.
  • JSON a model or a tool produced is refused beyond 512 levels of nesting. A reply, tool-call arguments or a structured-output document nested that deep is a Left (or a malformed tool call), not a StackOverflowError.
  • Caches. CacheKeyGenerator.sha256 now length-prefixes its parts, and the key function of CachedEmbeddingClient takes (text, modelName, purpose: InputPurpose); the default is CacheKeyGenerator.embeddingKey. A custom key function gains the parameter and should include it in the key. Vectors stored under the old keys in a persistent EmbeddingCache are not found and are re-embedded on first use. CacheConfig.create refuses a NaN similarity threshold.
  • Logged request and response bodies are redacted in more shapes (credentials under compound keys, in JSON inside a string, in arrays and objects, in key=value pairs), and a provider’s error body is now redacted before it is cut short for a log line or an error message, so a key straddling the cut no longer leaks as a fragment. Proxy-Authorization, Cookie, Set-Cookie and X-Amz-Security-Token are now sensitive in every shape, and a cookie header’s whole value is replaced; cookie is matched as a whole word, so cookie_policy is kept (#1686). Values in backslash-escaped quotes, JSON-escaped &, = and ? (\u0026, \u003d, \u003f), and key=value logs serialised into JSON strings are read too. Logs and error messages that showed a session cookie or such a value now show [REDACTED]; nothing else to change. See provider exchange logging for what is and is not redacted.
  • Llm4sHttpClient returns Result and takes FiniteDuration.
  • Removed, not deprecated. Before the compatibility baseline, dead, deprecated and accidental public API was deleted: for example Result.fromTry (use t.toResult), ToolBuilder#build() (use buildSafe()), ToolRegistry#getToolDefinitionsSafe (use getOpenAITools()), ReliableProviders and ReliabilitySyntax. The passes are listed, with their replacements, in the Migration Guide.

8. Java, Kotlin, Spring, cats-effect and ZIO

llm4s-java-api, llm4s-spring-boot-starter, llm4s-effect and llm4s-zio are new artifacts, and a Kotlin coroutine API is in the repository (Experimental). None of them was in 0.4.1, so there is nothing to migrate from it. If you use AgentIO or AgentZ, see Agent users.

If you built against a snapshot of main, these changed before the release:

  • Agent results are Java types. JAgent.run, continueConversation, resume, recover, AgentStream.await() and AgentStreamListener.onComplete return a JAgentResult, not the Scala AgentResult: answer() is an Optional<String>, messages() a java.util.List<JMessage>, usage() a JUsageSummary, and status().kind() the Java enum AgentStatusKind. The Kotlin AgentKt functions and AgentStreamItem.Done take or carry a JAgentResult too. Replace import org.llm4s.agent.AgentResult with org.llm4s.javaapi.JAgentResult; the CHANGELOG entry for #1393 has the full mapping.
  • Suspended turns. JAgent.pending(result), resume(threadId, answers) and recover(threadId) (and the Kotlin AgentKt.pending, resume and recover) handle a turn waiting for approval; see Suspended turns from Java and Kotlin (#1392). JAgent.pending(result) is now a shortcut for result.status().pending().
  • Kotlin cancellation. Cancelling AgentKt.run or continueConversation (a cancelled scope, withTimeout) now cancels the turn, as resume, recover and the streams do, and returns once it has ended, leaving the thread for recover; it used to stop only the wait, while the turn kept running and the thread stayed busy.
  • Interrupts. The Java facade never throws InterruptedException: an interrupted call is a CancelledError result, with the thread’s interrupt flag set. Test the result rather than catching the exception.
  • An interrupted JAgent call cancels its turn. The blocking JAgent.run, continueConversation, resume and recover used to stop only the wait and leave the turn running; they now cancel it and return once it has ended, leaving the thread for recover, as the Scala calls and Kotlin do (see Interrupting a blocking call cancels its turn). A turn that had already begun committing returns its real outcome with the interrupt flag set. AgentStream.cancel() still cancels a streamed turn without interrupting any thread, and now waits up to 5 seconds for it to end, as the blocking calls do, instead of without a bound.
  • Forgetting threads. Kotlin AgentKt.run(query), like Scala Agent.run(query) and Java JAgent.run(query), forgets the random thread of a turn that failed or was cancelled; it used to stay in the agent’s runtime. The new JAgent.forget(threadId) forgets a conversation by its thread id, for one whose turn left no result (#1682, #1688).
  • Whole replies, options and typed errors. JLlmClient.completion(...) returns an LlmResult<JCompletion> with the content, model, tool calls, token usage and estimated cost; complete still returns the text. JCompletionOptions.builder() sets temperature, top-p, token limits, penalties and JReasoningEffort without scala.Option. LlmException.getKind() gives an LlmErrorKind, with isRecoverable(), getRetryAfter() and getStatusCode(): call isRecoverable() rather than infer it from the kind (#1487, #1488).
  • Embeddings. Llm4s.createDefaultEmbeddingClient() returns a JEmbeddingClient for the model EMBEDDING_MODEL names. LlmException.getKind() now reads an embedding provider’s EmbeddingError by its HTTP status (AUTHENTICATION, RATE_LIMIT, VALIDATION, SERVICE) where it was always OTHER (#1490). See the Java guide.
  • These modules need JDK 21, like the rest of llm4s.

9. Provider and tracing authors

  • A provider is a ProviderDescriptor in its own module, supplied by an Llm4sProviderModule declared in META-INF/services/org.llm4s.llmconnect.spi.Llm4sProviderModule. It must be a class with a public no-argument constructor, not an object: ServiceLoader instantiates the named class.
  • Adding a provider is adding a dependency. ProviderRegistry.default is ProviderRegistry.discover(), computed once. A broken services entry is recorded as a failure and does not take out the others.
  • ProviderKind became the opaque type ProviderId, and ProviderConfig is open to providers outside core. The fromValues factories (OpenAIConfig.fromValues, AzureConfig.fromValues and the others) return a Result instead of throwing.
  • Provider-specific keys are declared in ProviderConfigSpec.extras and read from NamedProviderConfig.extras.
  • A tracing backend implements org.llm4s.trace.spi.TracingBackend, declared in META-INF/services/org.llm4s.trace.spi.TracingBackend, with mode = TracingMode.Named("yourmode").
  • llm4s-provider-testkit (ProviderModuleChecks) is how a provider module proves its registration.

See writing a provider.


What did not change

  • Most package names, apart from the extractor consolidation and the UsageSummary / ModelUsage move above.
  • Scala 3 (3.7.1) and JDK 21.
  • Named provider sections under llm4s.providers, and Llm4sConfig.defaultProvider().
  • The tool API: ToolFunction, ToolRegistry, ToolBuilder and Schema stay in llm4s-core, so your own tools need nothing new.
  • Agent.run(query) as the way to run an agent.

Before the release

For the maintainer who cuts 0.5.0. This page was written against main on 2026-10-07, brought up to date on 2026-10-10 and again on 2026-10-11 (through #1806 and #1722), and has to be re-checked:

  1. Version. The snippets show 0.5.0 literally, on purpose (migration guides are exempt from scripts/check-doc-versions.sh). Change them if the number changes.
  2. Artifacts. Run sbt -error listPublishedArtifacts and scripts/verify-release.sh, and compare the output with the tables in Dependencies.
  3. Anything merged since. Read the CHANGELOG’s Unreleased section for entries added after this page, and add what a user has to act on.
  4. The tests. sbt "samples/testOnly org.llm4s.samples.upgrade.UpgradeGuideSpec" checks the provider table, explicit registration, the agent snippet, an interrupted run, the removed orchestration, tracing discovery, the renamed RAG and error APIs, the redirectSafeHeaders snippet and whole-word memory search. sbt "core/testOnly org.llm4s.llmconnect.spi.CoreOnlyProviderDiscoverySpec" checks the core-only provider boundary.
  5. Links. Run scripts/check-doc-links.sh and scripts/check-doc-support.sh.
  6. The install guide. Remove the “Not yet published” notes in docs/getting-started/installation.md, and update the g8 starter template, which still pins 0.4.0.