Migration Guide
Stage 1 migration: agent runtime
Not in a release yet (#1328, with #1329’s events and tracing, which restore the agent event stream #1328 removed). Agent now runs on GraphRuntime: the graph is the only agent loop, AgentState and the legacy loop are deleted, and nothing runs the old loop beside the new one. Tools, guardrails, handoffs and context pruning belong to the agent, set when you build it, and a conversation is carried by its ThreadId instead of by a value you pass back in. Design: docs/design/typed-agent-runtime-design.md §4.13. This note covers the whole of Stage 1 (#1326), five slices:
- #1327, kernel completion: below.
- #1328, the agent loop on
GraphRuntime: this section. - #1329, events and tracing: the bullets this section marks #1329.
- #1330, orchestration removed: below.
- #1331, cancellation for non-chat clients: below.
1
2
3
4
5
6
7
8
9
10
11
12
13
// before
val agent = new Agent(client)
val state = agent.run("query", tools, inputGuardrails = in, outputGuardrails = out, maxSteps = Some(10))
// after
val result = for {
agent <- Agent.builder("assistant", client)
.withTools(tools)
.withMiddleware(new GuardrailMiddleware(in, out))
.withMaxSteps(10)
.build()
result <- agent.run("query")
} yield result
new Agent(client).run(q, tools, ...)isAgent.builder(id, client)...build()thenrun(q).tools,systemMessage,completionOptions,maxStepsandhandoffsmove towithTools(aToolRegistry,MCPToolRegistryincluded, or aToolSet),withSystemPrompt,withCompletionOptions,withMaxStepsandwithHandoffs.build()returnsResult[Agent]and refuses clashing tool names, an invalid or duplicate handoff id, andmaxSteps < 1.runhas the overloadsrun(query, config),run(threadId, query),run(threadId, query, config)andrun(threadId, query, config, history);configis aRunConfig(budgets, deadline).- Per-run guardrails are middleware.
inputGuardrails/outputGuardrailsarguments become.withMiddleware(new GuardrailMiddleware(input, output))on the builder. A block is no longer an error: the run returnsRightwithAgentStatus.Blocked(guardrail, reason)and the thread stays usable (its checkpoint isFailed, whichrecoverhas nothing to continue and the nextrunaccepts). An input block stores nothing of that turn’s query (a new thread still keeps its importedhistory); an output block removes the whole turn - query, tool calls and results, answer, and any handoff made in it - soresult.messagesis the history from before the turn;usagekeeps the turn’s model calls. Another middleware’sbeforeAgent/afterAgentLeftends the run the same way but is returned as thatLeft, and a blank query - given, or produced by abeforeAgent- is aValidationErrorwith nothing stored. A guardrail that transforms (PIIMasker) now applies its transformation. Guardrails - and any run-boundary middleware (beforeAgent,afterAgent) - on the root agent guard the whole handoff family: they apply to every turn’s query and final answer, whichever agent is active, outside the active agent’s own. Give them to the root only; model and tool wrappers stay per agent. continueConversation(state, ...)iscontinueConversation(result, ...). It reads onlyresult.threadId;run(threadId, query)is the same call by thread. A turn on aSuspendedresult is refused, not layered on parked work.- Threads are kept until forgotten. An
AgentStatewas garbage once dropped; a thread lives in the agent’s runtime, and one-shotrunthreads stay in the runtime untilforget. Callagent.forget(threadId)(orGraphRuntime.deleteThread(threadId, config)) for a conversation you will not continue.Checkpointerimplementations gaindeleteThread(threadId). runMultiTurnwithcontextWindowConfigisrunMultiTurn(first, followUps, config)on an agent built withnew ContextWindowMiddleware(config). Pruning trims what is sent to the model and never what the thread stores; the current turn (latest user message onward) is never pruned - the strategy,Customincluded, runs only on the history before it, with the budget the turn leaves - the request always starts with a user message (so it may exceed the budget), and the system prompt is outside the budget.runMultiTurnstops at the first status that is notCompleted.-
AgentStatefields move toAgentResult.AgentStateNow conversationresult.messages(the thread’s full history, never a system prompt)statusresult.status(below)logsremoved; use withTracingand thegraph.*eventsusageSummaryresult.usage(accumulated over the thread)tools,systemMessage,completionOptions,availableHandoffs,initialQuerythe builder; activeAgentis on the result AgentStatuscases are replaced.CompleteisCompleted(answer);Failed(error)isLeft(GraphError...)(provider errors asGraphError.NodeFailed(cause); the thread stays recoverable withrecover);InProgressandWaitingForToolsare not exposed;HandoffRequestedis gone, because a handoff runs inside the same run. New:Blocked(guardrail, reason),StepLimitReachedandSuspended(approvals, questions), continued withresume(threadId, answers)andresult.approve/reject/edit/reply.AgentContextis removed.tracingbecomesAgent.builder(...).withTracing(tracing), which emits the runtime’sgraph.*custom events for each run;debugandtraceLogPathare gone.TraceEvent.AgentStateUpdatedis no longer emitted by the agent; #1329 then replaced it withTraceEvent.AgentRunEnded.Handoff(agent)isHandoff(id, builder), andtransferSystemMessageis removed. The target is anAgentBuilder, not a built agent, and its id must equal the handoff id:Handoff.to("physics", physicsBuilder, "reason"). A target that must hand back to an agent already in the graph is named by id,Handoff.toId("triage", "reason"), so cycles compile. Each agent sends its own system prompt, which is never stored. A self-handoff is refused at build, and a handoff mixed with other tool calls in one message is refused at run time with an error result for every call.preserveContext = false(a parameter ofHandoff.to,Handoff.ofandHandoff.toId) sends the target the last user message before the transfer, then the transfer onwards.runStep,initializeSafe,runWithStrategyandToolExecutionStrategy(on the agent) are removed.agent.start(threadId, query)returns anAgentRun(threadId,runId,status,await(),cancel()) at once;runisstartthenawait;startRecoverandstartResumeare the same forrecoverandresume. The parallel tool calls of one message are bounded byRunBudgets.maxConcurrency.ToolExecutionStrategystays in core as aToolRegistryfeature; only the agent’s use of it is gone. Step-by-step observation isAgent.streamorAgentRun.subscribe(#1329, below).-
runWithEvents,continueConversationWithEvents,runCollectingEvents,AgentEventandAgentStreamingExecutorare removed (#1329 replaces them withAgent.stream, below).TraceEvent.AgentStateUpdatedis replaced byTraceEvent.AgentRunEnded, andAgentState#toTraceEventis gone.Before After agent.runWithEvents(query)(onEvent)agent.stream(threadId, query)(listener).flatMap(_.await())runCollectingEventscollect in the listener, or replay with GraphRuntime.subscribe(threadId, afterSeq = 0)AgentEvent.TextDelta(delta)AgentEvents.TextDelta(d)(d.text,d.attempt) withwithStreaming()ToolCallStarted/ToolCallCompleted/ToolCallFailedAgentEvents.ToolCallStarted,ToolCallResult(live),ToolExecuted(durable, with outcome)HandoffStarted/HandoffCompletedAgentEvents.HandedOffInputGuardrail*/OutputGuardrail*AgentEvents.GuardrailBlocked(on a block); the outcome onAgentStatus.BlockedAgentStarted/AgentCompleted/AgentFailed,StepStarted/StepCompletedkernel events: RunStarted,RunCompleted,RunFailed;ModelCallStarted/ModelCallCompletedTraceEvent.AgentStateUpdatedTraceEvent.AgentRunEndedcontext.progress(payload)context.progress(name, version, payload), or anEventTypeModelStep.next(messages, tools)next(messages, tools, call)Also:
AgentBuilder.withStreaming(),AgentRun.subscribe,AgentIO.stream*andAgentZ.stream*are new. Langfuse traces now use the run id as the trace id and the thread id as the session id.await()returns once the listener has returned from the run’s last event, so whatever it collected is complete.AgentRunEnded.usageis the run’s own usage, not the thread’s: a dashboard that summed the old cumulative figure per run double counted. A slowAgentIO.stream/AgentZ.streamconsumer loses live events and gets aStreamEvent.LiveGapwith their count; it no longer cancels the run. See the streaming guide. - Session files:
AgentState.saveToFile/loadFromFilebecome saved messages plushistory. Saveresult.messages(aVector[Message]has aReadWriter) from a completed run, and lateragent.run(newThreadId, query, RunConfig(), history = saved)on a thread that does not exist.historyis refused on an existing thread, with a system message in it (prompts belong to the agents), or if it ends mid tool call.ConversationPersistenceExampleshows it. Old session JSON files are not read. AgentIOandAgentZwrap the newAgent.LLMClientIO.agent(id)(configure)andLLMClientZ.agent(id)(configure)take a function over theAgentBuilder(tools now belong to the agent);run,continueConversation,recoverandresumetake aRunConfigand return theAgentResult. Cancelling the fiber cancels the run. A thrown exception arrives asGraphError.NodeFailedcarrying the original.CodeWorker.executeTasklosestraceLogPathand returnsResult[AgentResult];WorkspaceSettings.traceLogPathandWORKSPACE_TRACE_LOGare removed.AssistantAgentbuilds its agent once and keeps a thread id;SessionStateholds the thread id and the last result, and itsconsoleConfigparameter is gone.- Graph loop API.
ToolLoop.build(id, version, root, agents: Vector[LoopAgent])builds a family of agents (the earlierToolLoop.build(id, version, model, tools, middleware)is gone;LoopAgent(id, model, tools)withwithSystemPrompt,withMaxSteps,withMiddlewareandwithHandoffsreplaces its arguments), andModelStep.nextreturns aCompletion, so awrapModelCallmiddleware’snextreturnsResult[Completion]. - Errors. Provider, tool and middleware failures are
Left(GraphError...); the error content a tool failure gives the model is{"error": ...}. A guardrail block, the step limit and a suspension areRight. - Samples.
AsyncToolAgentExampleis deleted;StreamingAgentExample,StreamingWithToolsExampleandEventCollectionExampleare rewritten onAgent.stream(#1329); the other agent samples use the builder.
Kernel completion (#1327)
GraphBuilder.implement, node and resumeNode take retry: RetryPolicy and cache: Option[CachePolicy]. Both default to off, so a graph that sets neither runs as before. CompiledGraph.toMermaid is new. These are source breaks, with no shims:
ToolContext,GraphError.ToolFailed,ModelRequestandToolCallRequesthave a private constructor and no publiccopy. Build them withX(...)and change them withwithY(...).- Tool call ids and names are typed.
ToolContext.toolCallId,GraphError.ToolFailed.tooland.toolCallIdare the opaqueToolCallIdandToolNametypes: read the string with.value, and make an id withToolCallId(call.id).ToolCallRequestgainstoolCallIdandtoolName. recoverre-runs a failed task under the node’s full retry policy. It used to give the task one more try.
Cancellation for non-chat clients and MCP tool hints (#1331)
Embedding, reranker, MCP, image clients and the Whisper and Tacotron2 speech engines now return Left(CancelledError) when interrupted, as the cloud speech clients already did, with the thread’s interrupt flag still set, and never retry it. Before, they reported an error of their own, and some reported a cancelled call as a success. Match CancelledError where you matched those errors: an embedding provider’s embed can now return Left(CancelledError) where it returned only an EmbeddingError, so a match that narrows its Left to EmbeddingError needs a case for it. The CHANGELOG entry for #1331 lists every client’s change. Source breaks, with no shims:
- Image generation errors are
LLMErrors, with renamed cases: see Image generation errors areLLMErrors. MCPTransportImpl.sendRequest,sendNotification,MCPClient.initializeandgetToolsreturnResultinstead ofEither[String, _]: read the old string aserror.message.ToolHintsmoves tollm4s-core: importorg.llm4s.toolapi.ToolHints, notorg.llm4s.agent.graph.tool.ToolHints.
llm4s-mcp also reads ToolHints from a server’s tool annotations (MCPClient.getToolHints, MCPToolRegistry.toolHints(name)), but only for a server configured with trustAnnotations = true on its MCPServerConfig (default false). Any other server reports no hints, so ToolHints.default (approval required) applies.
Orchestration removed (#1330)
org.llm4s.agent.orchestration is deleted: PlanRunner, Plan, Node, Edge, TypedAgent, Policies, OrchestrationError and CancellationToken, with org.llm4s.types.PlanId and org.llm4s.types.AgentId from llm4s-core (the agent’s id is org.llm4s.agent.AgentId). PlanRunner passed Map[String, Any] between nodes and cast each node to TypedAgent[Any, Any]. A typed graph does the same job with checked handles, checkpoints and recovery. The multi-agent graph recipe is a worked replacement.
| Removed | Use instead |
|---|---|
TypedAgent[I, O], TypedAgent.fromFunction and the other factories |
a GraphNode[I] given to GraphBuilder.node; call an Agent inside the node for an LLM step |
Node, Edge, Plan, Plan.builder |
GraphBuilder.node / edge / staticJoin / dynamicJoin, then compile(entry)(output) |
PlanRunner.execute(plan, inputs, token) |
GraphRuntime.start(threadId, graph, input).flatMap(_.await()) |
PlanRunner(maxConcurrentNodes) |
RunConfig with RunBudgets(maxConcurrency = n) |
Policies.withRetry |
retry = RetryPolicy(...) on GraphBuilder.node / implement |
Policies.withTimeout |
RunBudgets.withTimeout (the whole run); a node bounds its own calls |
Policies.withFallback |
ordinary Result code in the node (primary.orElse(fallback)) |
OrchestrationError |
GraphError |
CancellationToken |
RunHandle.cancel() / AgentRun.cancel(), or interrupting the calling thread |
org.llm4s.types.PlanId |
RunId |
org.llm4s.types.AgentId |
org.llm4s.agent.AgentId |
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
// before
val plan = Plan.builder.addNode(research).addNode(summary).addEdge(Edge("e", research, summary)).build
val result = PlanRunner().execute(plan, Map("research" -> question), token) // Future[Result[Map[String, Any]]]
// after
val b = GraphBuilder("research", "v1")
val findings = StateKey.replace[String]("findings", "")
val digest = StateKey.replace[String]("digest", "")
val summary = b.node[Unit]("summary", writes = Set(digest)) { (_, state, _) =>
NodeResult.fromResult(for {
f <- state.get(findings)
turn <- summariser.run(s"Summarise: $f")
text <- turn.answer.toRight(ValidationError("summary", "no answer"))
} yield Command.empty.update(digest, text))
}
val research = b.node[String]("research", writes = Set(findings)) { (q, _, _) =>
NodeResult.fromResult(
researcher.run(q)
.flatMap(_.answer.toRight(ValidationError("research", "no answer")))
.map(f => Command.empty.update(findings, f).goto(summary))
)
}
val handle = b.compile(research)(_.get(digest)).flatMap(GraphRuntime.inMemory().start(ThreadId("t-1"), _, question))
handle.foreach(_.cancel()) // instead of token.cancel()
Agent.run, continueConversation, runMultiTurn, recover and resume now cancel their turn when the calling thread is interrupted, and return once it has ended (waiting up to 5 seconds for the turn to end), so recover can follow at once; a caller already interrupted starts no turn. Cancelling a graph run therefore also cancels the agent turns its nodes are waiting on. Before, the turn kept running after run returned Left(CancelledError). A caller that wants the turn to outlive an interrupt uses start, startRecover or startResume, and awaits the AgentRun itself. With tracing, the cancelled turn’s trace is complete when the call returns. A turn that had already begun committing its outcome when the interrupt came cannot be cancelled: the call returns that outcome (Right, Completed or Suspended), with the interrupt flag still set - test the flag, not only the result, if an interrupt must stop your own code. run(query), whose random thread id a Left does not carry, forgets the thread of a turn that failed or was cancelled once it has ended; name the thread (run(threadId, query)) to recover such a turn. The Java facade’s JAgent.run, continueConversation, resume and recover go through these calls, so an interrupted Java caller cancels its turn too; they used to stop only the wait.
Agent middleware
Not in a release yet (#1279). The graph tool loop
takes a stack of AgentMiddleware in place of a ToolCallPolicy: one ordered extension point with
beforeAgent, afterAgent, wrapModelCall and wrapToolCall hooks, for approval, guardrails,
logging, retry and rate limits. The graph runtime is Experimental, so these are source breaks with
no shims. Design: docs/design/typed-agent-runtime-design.md §4.8. (The ToolLoop.build signature shown below was generalised by #1328: see Stage 1 migration.)
-
ToolCallPolicyandPolicyDecisionare removed, andToolLoop.buildlosespolicy. It takesmiddleware: Seq[AgentMiddleware] = Nilinstead. A policy becomes anAgentMiddlewareoverridingwrapToolCall:Allowisnext(),Deny(reason)isToolOutcome.Error(s"Denied: $reason"), andRequireApproval(reason)isif context.approved then next() else ToolOutcome.NeedsApproval(reason)- or useApprovalMiddleware:1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20
// before val policy: ToolCallPolicy = call => call.name match case "drop_table" => PolicyDecision.Deny("never allowed") case "deploy" => PolicyDecision.RequireApproval("deploys are irreversible") case _ => PolicyDecision.Allow ToolLoop.build("assistant", "v1", model, tools, policy) // after val guard = new AgentMiddleware: val id = MiddlewareId("guard") override def wrapToolCall(request: ToolCallRequest, context: ToolContext)(next: () => ToolOutcome): ToolOutcome = request.call.name match case "drop_table" => ToolOutcome.Error("Denied: never allowed") case "deploy" if !context.approved => ToolOutcome.NeedsApproval("deploys are irreversible") case _ => next() ToolLoop.build("assistant", "v1", model, tools, Seq(guard)) // or, for approval alone ToolLoop.build("assistant", "v1", model, tools, Seq(ApprovalMiddleware.unlessReadOnly))
ApprovalSource.PolicybecomesApprovalSource.Middleware(id), naming the middleware that asked; a tool’s own request is stillApprovalSource.Tool.Approvenow re-runs the chain. It used to skip the policy and run the tool; it now runs the whole middleware chain again, from the outermost wrapper, withToolContext.approved = true, asEditdoes with the new arguments. A deny rule that depends only on the call refuses the same calls as before; a wrapper that asks for approval must pass whencontext.approvedis set, or the call becomes an error result (asked for approval again).- Guardrails in the graph loop apply their transformations.
GuardrailMiddleware(input, output)runs input guardrails inbeforeAgentand output guardrails inafterAgent, each on the value the previous one returned, so a guardrail that rewrites its input (PIIMasker, say) changes what the model sees and what the run answers. The legacyAgentvalidates every guardrail against the original value and keeps it, so it never applied a transformation; it is unchanged. ABlockfails the run with the errorCompositeGuardrail.allreports.
Agent tool contract and handoff ids
Not in a release yet (#1278). The graph tool loop
runs AgentTools, whose arguments are validated against their schema (rendered non-strict, so
optional fields may be omitted) before any middleware or the tool runs, and legacy handoffs take an explicit id. The graph runtime is
Experimental, so these are source breaks with no shims. Design:
docs/design/typed-agent-runtime-design.md §4.7.
(Since #1328, Handoff targets are builders and ToolLoop.build takes a family of LoopAgents: see
Stage 1 migration.)
-
LoopToolisAgentTool[A]. A tool now has anAgentToolSpec[A]- name, description and a coreSchemaDefinition[A], with aReadWriter[A]for the arguments - and receives them decoded, with aToolContext(run, call id, thread state,approved):1 2 3 4 5 6 7 8
// before val search = LoopTool("search")((call, approved) => ToolOutcome.Completed(run(call.arguments))) // after final case class SearchArgs(query: String) derives ReadWriter val schema = Schema.`object`[SearchArgs]("Search arguments").withRequiredField("query", Schema.string("Query")) val spec = AgentToolSpec[SearchArgs]("search", "Search the index", schema) val search = AgentTool(spec)((args, context) => ToolOutcome.Success(ujson.Str(run(args.query))))
LoopTool.fromToolFunction(f)becomesAgentTool.fromToolFunction(f); its arguments are now validated against the function’s schema too. toolloop.ToolOutcomeistool.ToolOutcome.Completed(content: String)becomesSuccess(content: ujson.Value, update)(aujson.Stris recorded as the string itself);Failed(message)becomesError(message);NeedsApprovalis unchanged.AskandFatalare new:Fatal(error)fails the run withGraphError.ToolFailed, reported insideGraphError.NodeFailed, and the run is recoverable. A thrown exception is still an error result.- Approval now arrives in the context.
LoopTool’sapprovedparameter iscontext.approved. - State updates are declared. A tool’s
Successupdate may touch only the keys in itswrites; anything else fails the run.ToolLoop.buildrefuses a tool that declaresToolLoop.resultsorMessages.key. ToolLoop.buildtakes aToolSet. Build it withToolSet.of(tools*), which returnsLeft(ValidationError)for an invalid name ([a-zA-Z0-9_-]{1,64}), a duplicate name, an argument schema that is not an object, or a schema keyword the validator cannot check.AgentToolSpec.applythrowsIllegalArgumentExceptionfor an invalid name.ModelStep.nexttakes the tool set.next(messages)becomesnext(messages, tools), andModelStep.fromClient(client, options)replacesoptions.toolswithtools.toolFunctions, so pass the tools toToolLoop.build, not inCompletionOptions.- Invalid arguments never reach the tool. Arguments that break the schema, fail to decode or
fail
withValidationbecome the error resultInvalid arguments for '<tool>': .... Edited approval arguments are checked again, and a middleware’s tool wrapper can still deny them. -
Handoffs take an id.
Handoff(agent, ...)becomesHandoff(id, agent, ...), andHandoff.to(agent)/Handoff.to(agent, reason)becomeHandoff.to(id, agent)/Handoff.to(id, agent, reason), which throwIllegalArgumentExceptionfor an invalid id;Handoff.of(id, agent, reason)returns aResult. An id matches[a-zA-Z0-9_-]{1,52}and is unique within one run’s handoffs; an agent run given an invalid or duplicate id fails withValidationErrorbefore any model call. The handoff tool ishandoff_to_<id>rather thanhandoff_to_agent_<hash>, so a conversation stored with the old name does not match a handoff.1 2 3 4 5
// before agent.run(query, tools, handoffs = Seq(Handoff.to(physicsAgent, "Physics expertise required"))) // after agent.run(query, tools, handoffs = Seq(Handoff.to("physics", physicsAgent, "Physics expertise required")))
Run API for graph runs
Not in a release yet (#1277). GraphRuntime.start,
recover and resume return a RunHandle once the thread is claimed, and the run executes on a
thread the runtime owns. The graph runtime is Experimental, so these are source breaks with no
shims. Design: docs/design/typed-agent-runtime-design.md §4.6.
NodeContextisRunContext.context.taskId,context.nodeIdandcontext.superstepbecomecontext.position.taskId,.nodeIdand.superstep;emitandprogressare unchanged, andisCancelledis new. Clients and tools stay captured by node closures; resolve per-tenant or per-thread resources fromcontext.config.tenantIdorcontext.position.- Limits are per run.
compile(entry, maxSupersteps)becomescompile(entry), andToolLoop.builddropsmaxSupersteps. PassRunConfig(budgets = RunBudgets(maxSupersteps = ..., timeout = ..., maxConcurrency = ...))to each run;RunBudgets.ofvalidates untrusted values as aResult, whileapplyand thewith*setters throwIllegalArgumentExceptionfor a non-positive value. CompiledGraph.run(input)is removed. UseGraphRuntime.inMemory().start(threadId, graph, input).flatMap(_.await()).step(execution)isstep(threadId, execution, config). It runs at mostconfig.budgets.maxConcurrencytasks at a time and checks no superstep limit.start/recover/resume(..., runId, durability)become(..., config, durability)and returnResult[RunHandle[O]]; the run id isconfig.runId. Callhandle.await()for theRunResult. Cancel withhandle.cancel(): interrupting the caller no longer cancels the run, and interrupting a thread blocked inawaitreturnsLeft(CancelledError)while the run continues. A cancel that arrives once the run has begun committing its outcome - completed, suspended or failed - is ignored, and the run ends with that outcome.recoverandresumetake the thread first, asstartdoes.recover(graph, threadId, ...)becomesrecover(threadId, graph, ...), andresume(graph, threadId, answers, ...)becomesresume(threadId, graph, answers, ...). Both arguments have different types, so the compiler flags every call to swap.- Subscribers have their own thread.
subscribegainscapacity(default 1024, at least 2). Listeners run on the subscription’s dispatcher thread, not a task or committing thread. A listener that falls behind by more thancapacitydurable events is disconnected; a listener that throws is disconnected rather than ignored.StreamEventgainsLiveGap(dropped)andDisconnected(lastSeq, reason); resubscribe withafterSeq = lastSeqto continue without a gap (TracingSubscriber.attachincluded: a lagging tracer is not re-attached for you). A subscription belongs to the thread, not one run, and keeps a parked dispatcher thread untilcancel(); cancel subscriptions you no longer need. - Tenants are checked. A run whose
RunConfig.tenantIddiffers from the one on the thread’s latest checkpoint is refused withGraphError.TenantMismatch(threadId, requested), which names only the caller’s tenant, never the owner’s;NoneandSomediffer. The check comes first: a wrong-tenant call getsTenantMismatchrather thanIncompleteRun,PendingInterrupts,NothingToRecover,NotSuspendedorThreadBusy. Checkpoints move to format 3, and earlier checkpoints read as having no tenant. GraphErroris classified per case. It extendsLLMErrorrather thanNonRecoverableError; every case is still aNonRecoverableErrorexcept the newDeadlineExceeded, which is aRecoverableError-recoverit with a new budget. Code that relied onGraphError <: NonRecoverableErrorshould match the case.- New cases break exhaustive matches.
GraphError.DeadlineExceeded,TenantMismatchandRunCrashed,RunEvent.RunTimedOut, andStreamEvent.LiveGapandDisconnectedare new.RunStartedis now a case class, and it,RunRecoveredandRunResumedgaintenantIdandprincipal; events in the old encoding still read.
Provider configs: build with apply, change with with*
Not in a release yet (#1388). OpenAIConfig,
AnthropicConfig and OllamaConfig follow the growth-prone pattern (pass 5 below), so a new field
no longer breaks Java and Kotlin callers. The constructor and copy are private:
- Scala:
OpenAIConfig(...)with the full field list, positional or named, still compiles. Replaceconfig.copy(baseUrl = url)withconfig.withBaseUrl(url); each field has a setter (withApiKey,withModel,withOrganization,withContextWindow, …), andOpenAIConfig’sOptionfields take the value or anOption. - Java and Kotlin: call the companion’s short
applyand chain setters instead of the constructor -OpenAIConfig.apply(apiKey, "gpt-4o").withOrganization("org-1"),AnthropicConfig.apply(apiKey, model),OllamaConfig.apply(model, baseUrl). It takes the default base URL and a context window from the model name;fromValuesstill consults the bundled model catalogue.
Image generation errors are LLMErrors
Not in a release yet (#1331). llm4s-image’s
generation clients return Either[LLMError, _], so an interrupted call can return
CancelledError, and ImageGenerationError extends LLMError. Source breaks, with no shims:
-
Five cases are renamed so they no longer share a name with an
org.llm4s.errortype. Both areLLMErrors, so amatchthat imported the wrong one would compile and never fire:Before ( org.llm4s.imagegeneration)After AuthenticationErrorImageAuthenticationErrorRateLimitErrorImageRateLimitErrorServiceErrorImageServiceErrorValidationErrorImageValidationErrorUnknownErrorImageUnknownError ImageServiceError(message, statusCode): the second field wascode: Int, which clashed withLLMError.code: Option[String]. It is a sealed type now (TransientImageServiceErrorfor0,408,429and5xx,RejectedImageServiceErrorotherwise); build and match it throughImageServiceError(message, status)as before.- A match on a client’s result needs a case for other
LLMErrors,CancelledErroramong them.
Cancellation by interrupt
Not in a release yet (#1270). An interrupted call
now returns Left(CancelledError) with the thread’s interrupt flag still set. Code that matched
ExecutionError, TimeoutError or SimpleError to detect an interrupted call should match
CancelledError instead, and ToolCallError.Cancelled for a tool call. CancelledError is
non-recoverable and is never retried. Chat clients no longer throw InterruptedException or report
UnknownError for an interrupt.
- New cases break exhaustive matches.
ErrorKind.Cancelled(metric labelcancelled),RunEvent.RunCancelledandToolCallError.Cancelledare new; amatchoverErrorKind,RunEventorToolCallErrorneeds a case for each. - Mid-stream,
NetworkErrorbecomesCancelledError. A stream read cut off by an interrupt used to be a recoverableNetworkError; code that retried it should stop onCancelledError. DefaultErrorMapperandTry(...).toResultclassify cancellations. An exception mapped while the thread is interrupted, or caused byInterruptedExceptionorClosedByInterruptException, becomesCancelledErrorrather thanUnknownErrororNetworkError. A bareInterruptedIOException- OkHttp’s call timeout is one - stays a timeout. Mapping never sets the interrupt flag; code that catches anInterruptedExceptionitself must restore it.ReliableClientreturnsCancelledError, not the localRateLimitError, when interrupted while waiting for a local rate-limit token.- Graph supersteps run concurrently. The default executor ran a superstep’s tasks one after
another; it now runs them concurrently on virtual threads, at most 16 at a time. Node code must be
thread-safe, a task does not inherit the caller’s
ThreadLocalor MDC context, and inSyncdurability, durable events and live progress are delivered on the task threads, not the caller’s. Since #1277 they are delivered on each subscription’s dispatcher thread instead (see above).
Pre-baseline API cleanup, pass 8
Not in a release yet; continues pass 7 below. It closes the last gaps in the frozen modules’
public API found by checking the re-audit
against main.
-
llm4s-agent’s console UI is internal.ConsoleInterface,ConsoleConfig(and itsStyleConfig) andMessageTypeareprivate[assistant]:ConsoleConfig’s colours were fansiAttrsandMessageTypecarried a publiccats.Show, which would have frozen both libraries intollm4s-agent’s API.AssistantAgentloses itsconsoleConfigparameter and the four-argument compatibility constructor; construct it with named arguments:1
new AssistantAgent(client, tools, sessionDir = "./sessions", agentContext = AgentContext.Default)
SessionState.localDateTimeRW, a public implicit upickle codec forjava.time.LocalDateTimethat anyimport SessionState._picked up, isprivate[assistant].org.llm4s.llmconnect.utils.SimilarityUtilsisprivate[llm4s]: only core’s caching client andllm4s-raguse it. Applications needing cosine similarity can compute it directly.
Pre-baseline API cleanup, pass 7
Not in a release yet; continues pass 6 below, which typed the times a caller supplies. This pass
types the times the library reports: an elapsed time is a FiniteDuration, a point in time an
Instant, and the unit leaves the name. Every JSON, trace and wire format keeps its keys and
millisecond values (duration_ms, durationMs, executionTimeMs, …).
1
2
3
4
// before
case AgentCompleted(state, steps, durationMs, _) => println(s"Done in ${durationMs}ms")
// after
case AgentCompleted(state, steps, duration, _) => println(s"Done in ${duration.toMillis}ms")
Check string interpolation in particular: s"${duration}ms" still compiles, but prints
150 millisecondsms.
| Module | Before | After |
|---|---|---|
llm4s-core |
TraceEvent.ToolExecuted.duration: Long, RAGOperationCompleted.durationMs, ImageGenerationCompleted.durationMs, Tracing.traceRAGOperation(durationMs) |
duration: FiniteDuration |
ProviderExchange.durationMs: Long |
duration: FiniteDuration |
|
RateLimitError.requestsRemaining, RateLimitError.resetTime |
removed: no constructor could set them, so they were always None |
|
llm4s-agent |
AgentEvent.ToolCallCompleted.durationMs, AgentCompleted.durationMs, AgentEvent.toolCompleted/agentCompleted’s durationMs |
duration: FiniteDuration |
llm4s-agent-tools |
ShellResult.executionTimeMs, HTTPResult.responseTimeMs |
executionTime, responseTime (the tool’s JSON keeps the ...Ms keys) |
llm4s-rag |
TimingInfo.durationMs, durationSeconds, avgPerItemMs |
duration, avgPerItem: Option[FiniteDuration] |
ExperimentResult.totalTimeMs/totalTimeSeconds |
totalTime: FiniteDuration |
|
BenchmarkResults.startTime/endTime: Long (epoch ms), totalDurationMs/totalDurationSeconds |
Instants, totalDuration: FiniteDuration |
|
llm4s-image |
ServiceStatus.averageGenerationTime: Option[Long] (ms) |
Option[FiniteDuration] |
llm4s-speech |
Transcription.processingTimeMs: Option[Long] |
processingTime: Option[FiniteDuration] |
| workspace | ExecuteCommandResponse.durationMs, CommandCompletedMessage.durationMs |
duration: FiniteDuration (the protocol keeps durationMs) |
Pre-baseline API cleanup, pass 6
Not in a release yet; continues pass 5 below. A time the caller supplies is typed: a duration is
a scala.concurrent.duration.FiniteDuration and a point in time a java.time.Instant, never a
raw Int/Long whose unit lives in its name or its Scaladoc. Names lose their unit suffix
(timeoutMs becomes timeout). Defaults are unchanged, and so are the wire formats: HOCON keys,
JSON sent to Exa and the workspace runner, and HTTP headers.
1
2
3
4
5
6
7
8
import scala.concurrent.duration.*
// before
Left(RateLimitError("openai", 1000L)) // milliseconds, documented as seconds
CrawlerConfig(delayMs = 500, timeoutMs = 30000)
// after
Left(RateLimitError("openai", 1.second))
CrawlerConfig(delay = 500.millis, timeout = 30.seconds)
llm4s-core errors, retry and reliability
RecoverableError.retryDelay,RateLimitError.retryAfter/retryDelayandServiceError.retryDelayareOption[FiniteDuration];RateLimitError(provider, retryAfter)takes aFiniteDurationand theRateLimitError(message, retryAfter, provider)extractor yields one.RateLimitError.resetTimeis anOption[Instant]. The fallback delay is namedRateLimitError.DefaultRetryDelay(30 seconds, as before).TimeoutError.timeoutDuration,ReliabilityConfig.deadline/withDeadline,CircuitBreakerConfig.recoveryTimeout/withRecoveryTimeout, everyRetryPolicydelay (exponentialBackoff,linearBackoff,fixedDelay,custom’s function,delayFor’s result) andErrorRecovery’srecoveryTimeoutwereDurationand areFiniteDuration: an infinite delay or deadline could not be slept or added to a clock. Code passing5.secondsis unchanged.ReliableClientandErrorRecovery.CircuitBreakertakeclock: () => Instant(was milliseconds since the epoch);ReliableClient’ssleep,ErrorRecovery’s andLLMClientRetry’ssleepFntake aFiniteDuration(was aLongof milliseconds).ErrorRecovery.recoverWithBackoff’sbaseDelayis aFiniteDuration.
llm4s-agent
OrchestrationError.AgentTimeoutError carries timeout: FiniteDuration (was timeoutMs: Long).
Beta modules
| Module | Before | After |
|---|---|---|
llm4s-agent-tools |
ExaSearchToolConfig/BraveSearchToolConfig/DuckDuckGoSearchToolConfig timeoutMs: Int, ShellConfig.timeoutMs: Long, HttpConfig.timeoutMs: Int |
timeout |
ExaSearchToolConfig.livecrawlTimeout: Option[Int] (ms) |
Option[FiniteDuration], still sent to Exa in ms |
|
llm4s-rag |
CrawlerConfig.delayMs/timeoutMs, withDelay(ms)/withTimeout(ms) (also on WebCrawlerLoader), UrlLoader.timeoutMs/withTimeout(ms) |
delay/timeout, withDelay(FiniteDuration)/withTimeout(FiniteDuration) |
RobotsTxtParser.isAllowed/getRules timeoutMs, RobotsTxt.crawlDelay: Option[Int] (s) |
timeout, Option[FiniteDuration] (fractional Crawl-delay values are now kept) |
|
EvaluatorOptions.timeoutMs, ChunkingUtils windowSeconds/clipSeconds, RateLimitedLogger throttleSeconds |
timeout, window/clip, throttle |
|
HikariDefaults.CONNECTION_TIMEOUT_MS/IDLE_TIMEOUT_MS/MAX_LIFETIME_MS |
ConnectionTimeout/IdleTimeout/MaxLifetime |
|
llm4s-mcp |
MCPServerConfig/transport timeout: Duration, MCPToolRegistry cacheTTL: Duration, MCPServer.stop(delay: Int) |
FiniteDuration |
StdioTransportImpl startupTimeoutMs: Int |
startupTimeout |
|
llm4s-image |
ImageGenerationConfig.timeout: Int (ms), the imagegeneration.provider.HttpClient methods’ timeout: Int |
FiniteDuration |
OpenAI/Anthropic vision configs’ connectTimeoutSeconds/requestTimeoutSeconds |
connectTimeout/requestTimeout |
|
| workspace | executeCommand timeout: Option[Int] (s), WorkspaceSandboxConfig.defaultCommandTimeoutSeconds |
Option[FiniteDuration], defaultCommandTimeout; the JSON still carries whole seconds under the old keys |
Test doubles of imagegeneration.provider.HttpClient type the timeout as FiniteDuration; a
ScalaMock onCall lambda typed _: Int compiles but fails at runtime.
Pre-baseline API cleanup, pass 5
Not in a release yet; continues pass 4 below. It settles the provider-author SPI’s API quality before the binary-compatibility baseline.
Growth-prone data types: construct with named arguments, change with with*
A case class cannot gain a field without breaking binary compatibility - its constructor,
apply and copy all change - so the types the library will plausibly extend now have a private
constructor, a public companion apply carrying the defaults, and with* setters:
CompletionOptions, Completion, StreamedChunk, TokenUsage, ModelCapabilities,
ModelMetadata, ProviderConfigSpec, EmbeddingConfigSpec, ProviderFeatures,
NamedProviderConfig, ReliabilityConfig, CircuitBreakerConfig, RateLimitConfig,
ContextConfig.
Construction is unchanged. .copy(...) is no longer available outside the type:
1
2
3
4
5
// before
val opts = CompletionOptions(temperature = 0.2).copy(maxTokens = Some(500))
// after
val opts = CompletionOptions(temperature = 0.2).withMaxTokens(500)
Every field has a withX. An Option field’s setter takes either the value or an Option
(withMaxTokens(500), withMaxTokens(None)). Match these types by name (c.usage), not by
position: a positional pattern breaks when a field is added.
Llm4sHttpClient returns Result and takes FiniteDuration
Every request method (get, post, postBytes, postMultipart, put, delete, postRaw,
postStream) returns Result[...] and never throws for a transport failure: a timeout is a
TimeoutError, a connection or I/O failure a NetworkError, an invalid URL, header or timeout a
ValidationError, an interruption an ExecutionError (with the interrupt flag restored). A
non-2xx status is still a Right. getResult is removed - get is now a Result itself.
1
2
3
4
5
// before
val response = Try(http.post(url, headers, body, timeout = 120000)).toResult
// after
import scala.concurrent.duration.*
val response: Result[HttpResponse] = http.post(url, headers, body, timeout = 120.seconds)
HttpRawResponse and StreamingHttpResponse carry headers, and every response type has a
case-insensitive header(name). Pass the headers to HttpErrorMapper.mapHttpError(status, body,
provider, headers) so a 429’s Retry-After (seconds or an HTTP date) becomes the
RateLimitError’s delay; every built-in provider does.
Test doubles that implement the trait return Result[...] and take FiniteDuration; return a
Left to simulate a failure instead of throwing. With ScalaMock, .returns(resp) becomes
.returns(Right(resp)) and .throws(e) becomes .returns(Left(error)); type the timeout in
onCall/where lambdas as FiniteDuration - a lambda typed _: Int still compiles but fails at
runtime.
StreamingAccumulator
| Before | After |
|---|---|
new StreamingAccumulator() |
StreamingAccumulator.create() (the class is final) |
getCurrentContent, getCurrentThinking, getCurrentToolCalls |
currentContent, currentThinking, currentToolCalls |
snapshot(), AccumulatorSnapshot |
toCompletion(created) |
StreamingAccumulator.withInitialState(...) |
create(), then addChunk / updateTokens |
RequestTransformer and TransformationResult
TransformationResult.warningsis removed: nothing ever filled it.TransformationResult.transform(modelId, options, messages, transformer, dropUnsupported = true)- the required
transformernow comes before the defaulted flag.
- the required
RequestTransformer#getDisallowedParamsis nowdisallowedParams.
llm4s-provider-testkit
New, Beta. A provider module’s spec mixes in org.llm4s.testkit.ProviderModuleChecks; see
Writing a provider. In this repository,
CredentialsRoundTrip and LocalProviderTestServer moved from core’s test sources to
org.llm4s.testkit.
Internal now
Llm4sConfig.providerFrom(source) and apiKeySourcesFrom(source) took a pureconfig
ConfigSource, which is not part of llm4s’s API; they are private[llm4s]. Load configuration
with Llm4sConfig.provider(name) / defaultProvider() / apiKeySources().
Pre-baseline API cleanup, pass 4
Not in a release yet; continues pass 3 below. Vendor fields leave NamedProviderConfig.
Config files: no change. organization, endpoint, apiVersion, contextWindow and
reserveCompletion keep their names in llm4s.providers.<name> sections. What changed is who owns
them: each is now a provider-specific key that only its provider declares.
| Key | Declared by |
|---|---|
organization |
openai, requesty, openrouter |
endpoint (required), apiVersion (default V2025_01_01_PREVIEW) |
azure |
contextWindow, reserveCompletion |
openai-compatible |
In a section for any other provider these keys used to be read and silently ignored; they are now
reported once as unknown keys, with a warning, and dropped - delete them. Vertex AI’s deprecated
endpoint/organization aliases for project/location work exactly as before. A non-numeric
contextWindow/reserveCompletion is now reported when the section is resolved, naming the key
(llm4s.providers.<name>.contextWindow must be a positive whole number, got '...'), rather than as a
type error while reading the block.
Scala callers:
NamedProviderConfigno longer hasorganization,endpoint,apiVersion,contextWindoworreserveCompletion. Read the validated value from the section’s extras:config.organization→config.extra("organization")(orOpenAIConfig.OrganizationKey),config.endpoint→config.extra(AzureProvider.EndpointKey),config.apiVersion→config.extra(AzureProvider.ApiVersionKey),config.contextWindow→config.extra(OpenAICompatibleProvider.ContextWindowKey).map(_.toInt)(likewiseReserveCompletionKey). Values are strings.- Code constructing
NamedProviderConfig(...)drops those arguments; pass them inextras = Map("endpoint" -> ..., ...)if the descriptor needs them. Positional callsNamedProviderConfig(id, model, baseUrl, apiKey, org, endpoint, apiVersion)becomeNamedProviderConfig(id, model, baseUrl, apiKey). ProviderConfigSpec(requiresEndpoint = true, endpointDescription = "...")→ProviderConfigSpec(extras = Seq(ProviderConfigKey.required("endpoint", "..."))).ProviderConfigSpec.BuiltinKeysis nowprovider, model, baseUrl, apiKey, headers, andBuiltinAliasKeysisbaseUrl, apiKey- adeprecatedAliasesentry naming one of the moved keys still works, resolved from the section’s extras.ProviderModelListers.openAICompatibleis now inllm4s-openai-compatible(packageorg.llm4s.configunchanged): a provider module calling it adds that dependency. It no longer sendsOpenAI-Organizationfrom the section; passsectionHeaders = ProviderModelListers.openAIOrganizationHeader(and declareOpenAIConfig.OrganizationConfigKey) to keep it, or anyNamedProviderConfig => Map[String, String]to derive other headers.ProviderModelListerandDiscoveredModelstay inllm4s-core.
Pre-baseline API cleanup, pass 3
Not in a release yet; continues pass 2 below.
llmconnect.middleware, ReliableProviders and ReliabilitySyntax are removed
The middleware package had no users outside its own tests and duplicated caching and the
metrics every provider client already records. LLMClient is a trait, so a decorator is a class
that implements it and delegates. ReliableProviders.wrap and .withReliability(...) were
shorthand for one constructor:
1
2
3
4
5
6
7
// before
val reliable = ReliableProviders.wrap(client, "openai", config)
val other = client.withReliability("openai")
// after
val reliable = new ReliableClient(client, "openai", config)
val other = new ReliableClient(client, "openai", ReliabilityConfig.default)
ReliableClient’s companion factories (ReliableClient(client), ReliableClient(client, config),
ReliableClient(client, config, metrics), ReliableClient.withProviderName) go too: three of them
guessed the provider name from the client’s class name. Use the constructor. For a client built
from config, the name is providerConfig.providerId.asString.
Behaviour fix: ReliableClient now applies ReliabilityConfig.rateLimit itself, before every
attempt, retries included. Before, only ReliableProviders.wrap honoured it, so a
new ReliableClient(...) with rate limiting enabled was not rate limited.
OpenAI’s model rules leave RequestTransformer
RequestTransformer.default now applies only the registry’s capabilities. The o-series
constraints (no system message, no native streaming, temperature 1, no sampling penalties) and
the max_completion_tokens rule for o-series and gpt-5 moved into llm4s-openai, which is the
only client that needs them; the Anthropic and Gemini clients no longer apply them to models
named like OpenAI’s.
| Removed | Use instead |
|---|---|
RequestTransformer#requiresMaxCompletionTokens, TransformationResult.requiresMaxCompletionTokens |
none in core; it is an OpenAI wire parameter |
DefaultRequestTransformer (now package-private) |
RequestTransformer.default(service) or withOverrides(...) |
| (new) | RequestTransformer.adjusted(service)((modelId, capabilities) => ...), for a provider module’s own rules on top of the registry |
Provider-author SPI
The plumbing provider modules build on is a public, frozen SPI; see
Writing a provider. Two OpenAI-format helpers moved to
llm4s-openai-compatible, with unchanged packages: org.llm4s.llmconnect.model.ResponseFormatMapper
and org.llm4s.llmconnect.serialization.{ToolCallDeserializer, StandardToolCallDeserializer}.
ProviderResultOps (tapRight / tapLeft) is now private[llm4s].
Pre-baseline API cleanup, pass 2
Not in a release yet; continues pass 1 below.
| Removed | Use instead |
|---|---|
ToolRegistry#getToolDefinitionsSafe(provider) |
getOpenAITools(). It returned the same JSON for every provider it accepted and failed for the rest; every client takes this format |
CancellationToken#cancellationFuture, cachedCancellationFuture |
whenCancelled, a shared Future[Unit] that succeeds on cancel: token.whenCancelled.map(_ => Left(myError)) |
CancellationToken#throwIfCancelled(), orchestration.CancellationException |
if (token.isCancelled) Left(...) |
ManagedResource.fileInputStream, dataOutputStream, byteArrayInputStream |
ManagedResource.fromTry(() => Try(new ...), s => Try(s.close())), or scala.util.Using |
ManagedResource map / flatMap (ManagedResourceOps) |
none: they never released the underlying resource. Nest use calls instead |
Behaviour fix: a PlanRunner node cancelled while running now fails the plan with
OrchestrationError.PlanExecutionError("Node <id> cancelled", ...). It used to surface as a
NodeExecutionError wrapping CancellationException, because the cancellation future only ever
failed and the mapping to the cancelled error never ran.
These utilities moved to the one module that uses them. Package names are unchanged, so code that
already depends on that module needs nothing; code that used them through llm4s-core alone adds
the module:
| Utility | Now in |
|---|---|
org.llm4s.util.SqlIdentifier, org.llm4s.llmconnect.utils.ChunkingUtils |
llm4s-rag |
org.llm4s.resource.ManagedResource |
llm4s-speech |
org.llm4s.util.LiftToResult |
llm4s-observability |
Pre-baseline API cleanup, pass 1
Not in a release yet; from the
spine re-audit. 0.5.0 sets
the binary-compatibility baseline for llm4s-core and llm4s-agent, so public API that nothing
used, or that was already deprecated, is removed now rather than frozen.
| Removed | Use instead |
|---|---|
Result.fromTry(t) |
t.toResult (import org.llm4s.types.TryOps) or Safety.fromTry(t) |
error.isRecoverable |
LLMError.isRecoverable(error), or match on RecoverableError / NonRecoverableError |
LLMError.fromThrowable(t) |
t.toLLMError (org.llm4s.error.ThrowableOps.RichThrowable) |
ToolBuilder#build() |
buildSafe(), which returns Result[ToolFunction] |
ToolRegistry#getToolDefinitions(p) |
getToolDefinitionsSafe(p) |
LLMCompressor.compress(...), LLMCompressedConversation |
ContextManager.withDefaults(tokenCounter, Some(llmClient)) then manageContext(conversation, budget); no direct equivalent (see below) |
Agent overloads taking debug, tracing or traceLogPath (run, runStep, continueConversation, runMultiTurn, runWithEvents, continueConversationWithEvents, runCollectingEvents, runWithStrategy, continueConversationWithStrategy) |
the same method with context = AgentContext(debug = ..., tracing = ..., traceLogPath = ...) |
ContextConfig(..., enableRollingSummary = ..., ...), ContextConfig.legacy(...) |
drop the argument (nothing read it); ContextConfig(...) or ContextConfig.default.copy(...) |
Safety.sequenceV(xs) |
xs.traverse(_.toValidatedNec) with import cats.syntax.all.*, which keeps every error; or, given items and a validator, Result.validateAll(items)(validate). Not Result.sequence, which stops at the first error |
LLMError.llmErrorShow, error.show, error.display |
error.formatted |
org.llm4s.types aliases and wrappers the library never used (CompletionId, ToolName, ToolCallId, Url, MessageId, WorkspaceId, Json, Timeout, TokenCount, the MCP/image/audio/video/plugin/workflow/metrics types, …) |
the underlying type, e.g. String, ujson.Value, Long, Int |
ConnectionStatus, ProviderCapabilities, ClientHealth, StreamingOptions, RuntimeId, ModelId |
none; nothing used them |
LLMCompressor.compress summarised a whole conversation with one LLM call, optionally with a
custom prompt. Nothing replaces it one-for-one: LLMCompressor.squeezeDigest only shrinks messages
already marked [HISTORY_SUMMARY] and returns any other conversation unchanged, so do not swap
one call for the other. ContextManager.manageContext is the supported path - it builds the
digests, compresses them and trims to the budget (ContextConfig.enableLLMCompression gates the
LLM step); the custom prompt has no counterpart.
RateLimitedLogger, ProvidersConfigModel.RawNamedProviderSection / RawProvidersConfig,
agent.orchestration.MDCContext and assistant.ShowInstances are now package-private.
org.llm4s.types keeps Result, AsyncResult, TryOps / OptionOps / FutureOps, and the
newtypes the library’s APIs take: SessionId, TraceId, FilePath, DirectoryPath, AgentId,
PlanId (both removed later with orchestration, #1330), SemanticBlockId, ArtifactKey, ExternalizedContent, ContentSize,
HeadroomPercent and the TokenBudget, ContextWindowSize, ByteCount and
ExternalizationThreshold aliases.
Slice 7: llm4s-agent - the agent runtime leaves core
The second slice 7 carve (#1242, decision D4)
moves the agent runtime into a new module, llm4s-agent. Package names are unchanged, so no
import changes; code that uses the agent adds one dependency:
1
libraryDependencies += "org.llm4s" %% "llm4s-agent" % "<version>"
Moved to llm4s-agent |
Package |
|---|---|
Agent, AgentState, AgentStatus, AgentContext, Handoff, ContextWindowConfig, AgentTraceFormatter |
org.llm4s.agent |
Guardrails - the traits, CompositeGuardrail, the built-in validators, the LLM-as-judge and RAG guardrails, PII patterns |
org.llm4s.agent.guardrails.* |
| DAG orchestration | org.llm4s.agent.orchestration |
Streaming events (AgentEvent and its cases) |
org.llm4s.agent.streaming |
The console assistant: AssistantAgent, SessionManager, ConsoleInterface (Beta) |
org.llm4s.assistant |
Not moved: agent memory (org.llm4s.agent.memory) is already llm4s-memory, which does not
depend on llm4s-agent. UsageSummary and ModelUsage moved to org.llm4s.llmconnect.model in
llm4s-core just before this (see below). The ready-made tools are llm4s-agent-tools.
What stays in llm4s-core is what the agent is built on: LLMClient and LLMConnect, the
tool API (ToolFunction, ToolRegistry), and the tracing contract. An agent run is traced through
TraceEvent.AgentStateUpdated, built by AgentState#toTraceEvent, so tracing backends need
nothing from the agent module.
Other modules: llm4s-workspace-client now depends on llm4s-agent, for its codegen package.
fansi leaves llm4s-core with the assistant, its only user.
Slice 7: llm4s-agent-tools - the built-in tools leave core
The first slice 7 carve (#1242, decisions D2 and D3)
moves the ready-made tools into a new module, llm4s-agent-tools. Package names are unchanged,
so no import changes; code that uses any of them adds one dependency:
1
libraryDependencies += "org.llm4s" %% "llm4s-agent-tools" % "<version>"
Moved to llm4s-agent-tools |
Package |
|---|---|
BuiltinTools and every tool under it: core (DateTime, Calculator, UUID, JSON), filesystem, http, shell, search (Brave, DuckDuckGo, Exa) |
org.llm4s.toolapi.builtin.* |
WeatherTool |
org.llm4s.toolapi.tools |
ToolsConfigLoader (was private[config], now public), BraveSearchToolConfig, DuckDuckGoSearchToolConfig, ExaSearchToolConfig |
org.llm4s.config |
ToolsConfigKeys (new) |
org.llm4s.config |
What stays in llm4s-core is the tool API: ToolFunction, ToolRegistry, ToolBuilder,
Schema, the execution strategies and SafeParameterExtractor. Your own tools need nothing new.
llm4s-agent-tools depends on llm4s-core only - not on the agent runtime - so the tools work
with plain tool calling through ToolRegistry as well as with Agent.
Configuration is unchanged
The llm4s.tools.brave, .duckduckgo and .exa keys, their defaults, and the
BRAVE_SEARCH_API_KEY, BRAVE_SEARCH_COUNT, BRAVE_SEARCH_API_URL, BRAVE_SAFE_SEARCH,
DUCK_DUCK_GO_SEARCH_API_URL and EXA_* variables are the same. The block moved from core’s
reference.conf to the module’s, beside the code that reads it.
Source breaks
- The tools need
llm4s-agent-tools(table above). -
Llm4sConfig.loadBraveSearchTool(),loadDuckDuckGoSearchTool()andloadExaSearchTool()are removed. They returned config types that left core. The same methods are onToolsConfigLoader:1 2 3 4 5 6 7
import org.llm4s.config.ToolsConfigLoader // before val braveConfig = Llm4sConfig.loadBraveSearchTool() // after val braveConfig = ToolsConfigLoader.loadBraveSearchTool()
Each also takes a
ConfigSource, to read from one of your own. ConfigKeys.BRAVE_SEARCH_API_KEYisToolsConfigKeys.BRAVE_SEARCH_API_KEY, withToolsConfigKeys.EXA_API_KEYbeside it.
Slice 7: UsageSummary and ModelUsage move to org.llm4s.llmconnect.model
Not in a release yet; the preparation step of slice 7 (#1242, decision D1).
Slice 7 carves the agent runtime out of llm4s-core into llm4s-agent. UsageSummary and
ModelUsage were in org.llm4s.agent, but they depend only on TokenUsage, and
llm4s-observability’s CostTracker builds them too. Leaving them in the agent package would make
llm4s-observability depend on the agent runtime; moving the file to core while keeping its
package would split org.llm4s.agent across two jars. So they move to org.llm4s.llmconnect.model,
beside TokenUsage, and stay in llm4s-core.
This is the one change to an import in the slice; every other move keeps its package.
1
2
3
4
5
// before
import org.llm4s.agent.{ ModelUsage, UsageSummary }
// after
import org.llm4s.llmconnect.model.{ ModelUsage, UsageSummary }
Code that only reads AgentState.usageSummary or CostTracker.snapshot needs no change. The JSON
form is unchanged, so an AgentState saved before the move still loads.
llm4s no longer brings a logging backend
Not in a release yet; found by the slice 6 spine audit (#1133).
Every llm4s module used to declare logback-classic and log4j-to-slf4j as compile
dependencies, so they reached your classpath with llm4s. Libraries should leave that choice to the
application, and the published modules now declare only slf4j-api.
- If you already configure logging (your own logback, log4j2 via
log4j-slf4j2-impl, or another SLF4J 2 backend), nothing changes - except that llm4s no longer adds a second backend or a bridge that clashes withlog4j-core. -
If you relied on the logback llm4s brought, declare it yourself; without a backend, SLF4J logs nothing and prints one “no SLF4J providers were found” warning:
1
libraryDependencies += "ch.qos.logback" % "logback-classic" % "1.5.34"
- If you used Monocle, commons-io or fansi through llm4s, declare them yourself: Monocle and
commons-io are no longer dependencies of any llm4s module, and fansi comes only with
llm4s-core.
A failed read no longer deletes or clears indexed documents
Follow-up to #1236; not in a release yet.
LoadResult.Failure gains a fourth field, documentId: Option[String] = None: the id the
failed document is (or would be) indexed under. source stays the path, key or URL. RAG.sync
uses it to tell a document it could not read from one that has gone from the source.
- A pattern match on
LoadResult.Failureneeds a fourth argument:case LoadResult.Failure(source, error, recoverable)becomescase LoadResult.Failure(source, error, recoverable, documentId)(or_). Constructor calls,LoadResult.failure(source, error),.source,.errorand.recoverableare unchanged;LoadResult.failure(source, error, documentId = id)is new. synckeeps a document whose read fails. AFailurewith adocumentIdkeeps that document’s indexed version: it is not deleted, and not counted inSyncStats. Previously sync deleted it, because it had not been “seen”.SourceBackedLoader(S3 and every otherDocumentSource) andUrlLoadername the document on every per-document failure,FileLoaderwhen extraction fails. Read failures are logged at WARN.- A
Failurewith nodocumentIdmakessyncskip its deletion pass for that run, as it cannot tell which unlisted document failed. Adds and updates still apply.WebCrawlerLoaderleaves it out on purpose (the pages below a failed page went uncrawled), as doesFileLoaderfor a missing path. A customDocumentLoadershould setdocumentIdon per-document failures when it knows the id, or its failures will now hold back deletions. refreshreads the whole loader before it clears anything. AListingFailure- at any point, including a later S3 page - or, underfailFast, anyFailurereturns the error and leaves the index and registry untouched. Previously the index was cleared first and left empty or half-rebuilt. The cost is memory:refreshandrefreshAsyncnow hold every loaded document’s extracted text in memory before clearing; usesync, which streams, for a source too large for that. WithoutfailFast, a document that fails to read is still left out of the rebuilt index.DocumentLoaders.successesOnlystill drops per-document failures, so a sync through it still deletes a document whose read failed.
Slice 6: llm4s-observability-prometheus - Prometheus leaves core
The second slice 6 carve (#1133, decisions D3 and
D4) moves the Prometheus metrics backend into a new module, llm4s-observability-prometheus,
which carries the Prometheus client and HTTP server. With it gone, llm4s-core declares no
observability dependency. Package names are unchanged, so no import changes:
1
libraryDependencies += "org.llm4s" %% "llm4s-observability-prometheus" % "<version>"
Moved to llm4s-observability-prometheus |
Package |
|---|---|
PrometheusMetrics, PrometheusEndpoint |
org.llm4s.metrics |
MetricsConfigLoader (was private[config], now public) |
org.llm4s.config |
What stays in llm4s-core is the metrics contract: MetricsCollector (with noop and
compose), Outcome and ErrorKind. Every provider client, ReliableClient and the metrics and
rate-limiting middleware take a MetricsCollector, so none of them needs the new module; only the
code that builds a PrometheusMetrics does. It is kept apart from llm4s-observability so that
Prometheus does not reach every llm4s-rag user through that module.
Configuration is unchanged
The keys and their defaults are the same - llm4s.metrics.enabled (false),
llm4s.metrics.prometheus.enabled (true) and llm4s.metrics.prometheus.port (9090). They used
to be hard-coded in MetricsConfigLoader, because core’s reference.conf had no llm4s.metrics
block; that block now ships in the module’s reference.conf, beside the code that reads it.
Source break: Llm4sConfig.metrics() is removed
It returned a PrometheusEndpoint, so it could not stay in a core without Prometheus (D3).
MetricsConfigLoader replaces it, with the same result type:
1
2
3
4
5
6
7
import org.llm4s.config.MetricsConfigLoader
// before
val (metrics, endpoint) = Llm4sConfig.metrics().toOption.get
// after
val (metrics, endpoint) = MetricsConfigLoader.default().toOption.get
MetricsConfigLoader.load(source) reads from a ConfigSource of your own. Its source parameter
no longer defaults to ConfigSource.default; call default() for that.
A failed listing fails the sync
#1231; not in a release yet.
A document loader that cannot enumerate its documents - an S3 listing with no credentials, a
missing bucket or access denied, a DirectoryLoader path that does not exist - now yields the
new LoadResult.ListingFailure(source, error) instead of a LoadResult.Failure. RAG.sync,
ingest and refresh, and their async forms, return its error as the Left. Previously
sync returned Right with 0 documents and then deleted every document it had indexed from
that source, because it had seen none of them.
- Code that matches on
LoadResultneeds aListingFailurecase; an exhaustive match without one warns, and fails the build under-Werror. RAG.ingestno longer counts a listing error as one failed document inLoadStats: it returns it as theLeft, whether or notfailFastis set. Code that readstats.errorsfor"list-error"should handle theLeftinstead.SourceBackedLoadernames the source, not"list-error": theListingFailure’ssourceis theDocumentSource’sdescription, e.g.S3(s3://bucket/prefix).- A
syncthat fails part-way through a listing (a later S3 page) keeps the documents it synced before the failure and deletes nothing; re-running it once the source is reachable completes the sync.refreshcleared the index before reading the loader, so a listing failure left it empty - but said so; it now leaves the index untouched (see above).
A DocumentSource reports a listing error by yielding a Left from listDocuments(), as
before; SourceBackedLoader does the rest. A custom DocumentLoader whose enumeration fails
should return LoadResult.listingFailure(source, error).
Vendor credentials: a shared API key per provider
#1132, part of the modularisation programme
(#1126); not in a release yet. 0.4.1 and earlier
behave as before.
Credentials belong to a vendor, keyed by provider id; clients belong to a use. Each provider
module now binds its vendor’s conventional API-key variable to a shared key,
llm4s.credentials.<provider>.apiKey, in its own reference.conf. A client - a chat section, an
embeddings block, the Cohere reranker - that sets no apiKey of its own uses it. So with
OPENAI_API_KEY set, a chat section needs only provider and model, and OpenAI embeddings need
no extra line.
This is additive: configs that set apiKey keep working unchanged, because a client’s own key
always wins. Configs that failed for want of an apiKey line now load.
Resolution order
For a client of provider P:
- its own
apiKey:llm4s.providers.<name>.apiKey,llm4s.embeddings.<P>.apiKey, orllm4s.rerank.cohere.apiKey; - otherwise
llm4s.credentials.<P>.apiKey, where P is the canonical id - aprovider = "google"section usesllm4s.credentials.gemini; - otherwise a
ConfigurationErrornaming both places:apiKey: set OPENAI_API_KEY, or set apiKey under llm4s.providers.openai-main in application.conf.
The credentials block holds apiKey only: baseUrl, endpoint, model and apiVersion are
never defaulted from it. Where each key came from is logged at INFO - for example
llm4s.providers.openai-main: API key from llm4s.credentials.openai.apiKey - and the value never
is.
1
2
3
4
5
6
7
8
9
10
11
12
13
# before
openai-main {
provider = "openai"
model = "gpt-4o-mini"
apiKey = ${?OPENAI_API_KEY}
}
llm4s.embeddings.openai.apiKey = ${?OPENAI_API_KEY}
# after - with OPENAI_API_KEY exported
openai-main {
provider = "openai"
model = "gpt-4o-mini"
}
The old lines still work and may stay. A section for a second account keeps its own key:
apiKey = ${?OPENAI_BATCH_API_KEY}. If that variable is unset, the section now falls back to the
shared key instead of failing - see Explicit keys in production.
Variables bound
| Provider id | Variable | Module |
|---|---|---|
openai |
OPENAI_API_KEY |
llm4s-openai |
azure |
AZURE_OPENAI_API_KEY |
llm4s-openai |
requesty |
REQUESTY_API_KEY |
llm4s-openai |
anthropic |
ANTHROPIC_API_KEY |
llm4s-anthropic |
gemini |
GOOGLE_API_KEY, else GEMINI_API_KEY (Google’s SDK precedence) |
llm4s-gemini |
deepseek, zai, openrouter, mistral |
DEEPSEEK_API_KEY, ZAI_API_KEY, OPENROUTER_API_KEY, MISTRAL_API_KEY |
llm4s-openai-compatible |
cohere |
COHERE_API_KEY |
llm4s-openai-compatible and llm4s-rag |
voyage |
VOYAGE_API_KEY |
llm4s-voyage |
None for openai-compatible (it has no vendor), ollama (no key) or vertexai (OAuth2:
Application Default Credentials, or a service-account file named by apiKey).
Moved bindings
- Voyage:
llm4s.embeddings.voyage.apiKey = ${?VOYAGE_API_KEY}becamellm4s.credentials.voyage.apiKey = ${?VOYAGE_API_KEY}.VOYAGE_API_KEYworks as before, and an explicitllm4s.embeddings.voyage.apiKeystill wins. - Cohere reranker:
llm4s.rerank.cohere.apiKey = ${?COHERE_API_KEY}becamellm4s.credentials.cohere.apiKey = ${?COHERE_API_KEY}, shared with the Cohere chat provider, so oneCOHERE_API_KEYserves both. An explicitllm4s.rerank.cohere.apiKeystill wins. - Azure: if you exported
AZURE_API_KEYfor a section without its ownapiKey, rename it toAZURE_OPENAI_API_KEY- the openai SDK’s name, and the one now bound - or keep anapiKey = ${?AZURE_API_KEY}line in the section.
New: RerankerConfigLoader
llm4s-rag’s reference.conf has bound llm4s.rerank since #337, but no code read it.
org.llm4s.config.RerankerConfigLoader.load(source) / .default() now does, returning
Result[Option[RerankProviderConfig]] for RAG.build(..., resolveRerankerConfig = ...) or
RerankerFactory.fromConfig.
Explicit keys in production
llm4s-config-policy’s prod preset (ConfigPolicy.prodSafeDefaults) gains the rule
ownApiKey (ConfigPolicy.withOwnApiKeyRequired, checked by
ConfigPolicyEngine.checkApiKeySources and run by CheckPolicies): every chat section whose
provider requires a key must set its own apiKey, so a section meant for a second account cannot
silently bill the default one. A prod check of a config that relies on the shared key now fails;
add apiKey = ${?VAR} to each section. Llm4sConfig.apiKeySources() / apiKeySourcesFrom(source)
report each section’s ApiKeySource (Section(path) or Credentials(path)) for checks of your own.
Source breaks
These are deliberate removals ahead of the MiMa baseline:
EmbeddingConfigSpec.apiKeyPathis removed, with the loader code that resolved it. The sharedllm4s.credentials.<id>.apiKeyreplaces it for every provider: bind your vendor’s variable there in your module’sreference.confinstead of declaring a path.EmbeddingConfigSpec.apiKeyEnvis aSeq[String], not anOption[String]:apiKeyEnv = Some("X")becomesapiKeyEnv = Seq("X"). It lists the variables your module binds tollm4s.credentials.<id>.apiKey, highest precedence first, and is used only to word the missing-key error.ProviderConfigSpecgainedapiKeyEnv: Seq[String](last, with a default), andProviderConfigSpec.apiKeyAndDefaultBaseUrlan optional second parameter for it. Positional construction still compiles; a pattern match onProviderConfigSpecneeds one more binder.OpenAIConfigKeys.AZURE_API_KEYis nowAZURE_OPENAI_API_KEY, with the value"AZURE_OPENAI_API_KEY".- The missing-
apiKeymessages changed. Chat:apiKey: set <VAR>, or set apiKey under llm4s.providers.<name> in application.conf. Embeddings:Missing <id> embeddings apiKey: set <VAR>, or set apiKey under llm4s.embeddings.<id> in application.conf. Code matching the old text needs updating.
For provider authors: bind llm4s.credentials.<id>.apiKey = ${?<VENDOR>_API_KEY} in your module’s
reference.conf, using the name the vendor’s own SDK or docs use, and declare the same variable
as apiKeyEnv on your ProviderConfigSpec or EmbeddingConfigSpec. Your module’s round-trip spec
should prove the two agree.
Slice 6: llm4s-observability - Langfuse, the trace collector and CostTracker leave core
The slice 6 carve (#1133,
#1126) moves the tracing integrations that need
nothing beyond core into a new module, llm4s-observability. Package names are unchanged, so
no import changes; code that uses any of the types below adds one dependency:
1
libraryDependencies += "org.llm4s" %% "llm4s-observability" % "<version>"
Moved to llm4s-observability |
Package |
|---|---|
LangfuseTracing, LangfuseBatchSender, DefaultLangfuseBatchSender, LangfuseHttpApiCaller, LangfuseTracingBackend |
org.llm4s.trace |
TraceCollectorTracing and the rest of TraceCollector.scala |
org.llm4s.trace |
Trace, Span, SpanId, SpanKind, SpanStatus, SpanEvent, SpanValue, TraceModelJson |
org.llm4s.trace.model |
TraceStore, InMemoryTraceStore, TraceQuery |
org.llm4s.trace.store |
CostTracker |
org.llm4s.metrics |
LangfuseConfig (was in core) |
org.llm4s.llmconnect.config |
LangfuseConfigKeys, LangfuseConfigLoader (new) |
org.llm4s.config |
OpenTelemetryConfig moves, in the same package, to llm4s-observability-otel, beside the
backend that reads it.
What stays in llm4s-core is the tracing contract (decisions D1 to D5 in #1133): Tracing,
TraceEvent, TracingComposer, TracingMode, the TracingBackend SPI, NoOpTracing,
ConsoleTracing (the default mode, so a core-only application still traces), TracingSettings,
and MetricsCollector. llmconnect, the agent runtime and every provider module use only these,
so none of them needs the new module. llm4s-rag depends on it, for RAGASLangfuseObserver; it
carries no third-party dependency, so this adds nothing else to a RAG user’s classpath.
Configuration is unchanged
The keys and variables are the same: TRACING_MODE=langfuse, LANGFUSE_PUBLIC_KEY,
LANGFUSE_SECRET_KEY, LANGFUSE_URL, LANGFUSE_ENV, LANGFUSE_RELEASE, LANGFUSE_VERSION
under llm4s.tracing.langfuse.*, and TRACING_MODE=opentelemetry (or otel),
OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT and llm4s.tracing.opentelemetry.headers.
Only the module that binds them changed: the llm4s.tracing.langfuse block moved from core’s
reference.conf to llm4s-observability’s, and llm4s.tracing.opentelemetry to
llm4s-observability-otel’s, each with the code that reads it. Core now reads only
llm4s.tracing.mode, and hands the selected mode’s block to its backend as
TracingSettings.extras.
So an application that sets TRACING_MODE=langfuse adds llm4s-observability and changes no
config. Without it, Tracing.create logs an error naming the artifact and traces nothing -
exactly what TRACING_MODE=opentelemetry without llm4s-observability-otel already did:
1
2
Tracing mode 'langfuse' is configured but no TracingBackend for it is on the classpath.
Add the 'org.llm4s' %% 'llm4s-observability' dependency. Available modes: console, noop.
Behaviour change: Langfuse without keys does not start
TRACING_MODE=langfuse without LANGFUSE_PUBLIC_KEY or LANGFUSE_SECRET_KEY used to build a
LangfuseTracing that logged a warning and dropped every batch. The backend now refuses to start:
Tracing.fromSettings returns
1
2
Langfuse tracing is selected but llm4s.tracing.langfuse.publicKey (LANGFUSE_PUBLIC_KEY) and
llm4s.tracing.langfuse.secretKey (LANGFUSE_SECRET_KEY) are not set.
and Tracing.create logs that once and returns NoOpTracing. Either way nothing reaches
Langfuse, as before; the difference is one clear error instead of a warning per batch.
LangfuseTracing.from(config) built directly is unchanged.
Behaviour change: a selected block that is not an object is an error
TracingSettings.extras is the selected mode’s block, llm4s.tracing.<mode>. When that key was
present but was not an object - llm4s.tracing { mode = opentelemetry, opentelemetry =
"http://collector:4317" } - it was treated as absent, so the backend started on its defaults
(here, a collector on localhost) and the operator’s setting was silently ignored. Read failures
were swallowed the same way. Llm4sConfig.tracing() now returns a ConfigurationError naming
the path instead:
1
2
llm4s.tracing.opentelemetry must be an object, but is a string. It holds the settings for
tracing mode 'opentelemetry': write it as llm4s.tracing.opentelemetry { ... }.
A block that is absent (or null) is still an empty extras, and the backend applies its
defaults. Only the selected mode’s block is checked: a malformed block for a mode that is not
selected is still ignored.
Reading Langfuse settings yourself
TracingSettings no longer has a langfuse field, and extras holds only the selected mode’s
block. To read llm4s.tracing.langfuse whatever TRACING_MODE is - to combine Langfuse with
Console, or for RAGASLangfuseObserver - use the new loader:
1
2
3
4
5
6
7
8
9
import org.llm4s.config.LangfuseConfigLoader
// before
Llm4sConfig.tracing().map(settings => LangfuseTracing.from(settings.langfuse))
RAGASLangfuseObserver.fromTracingSettings(settings)
// after
LangfuseConfigLoader.default().map(LangfuseTracing.from)
LangfuseConfigLoader.default().map(RAGASLangfuseObserver.from)
Source breaks
These are taken before 0.5.0 sets the MiMa baseline, and none of them changes a package name.
- Langfuse, the collector, its model and store, and
CostTrackerneedllm4s-observability(table above). TracingMode.LangfuseandTracingMode.OpenTelemetryare removed. Core keeps a case only for what it builds itself,ConsoleandNoOp; the others areTracingMode.Named("langfuse")andTracingMode.Named("opentelemetry"), whichTracingMode.fromStringreturns for"langfuse"and for"opentelemetry"/"otel".LangfuseConfig.ModeandOpenTelemetryConfig.Modename them. Amatchon the old case objects matches onTracingMode.Named("langfuse")instead.TracingSettingsisTracingSettings(mode, extras). Thelangfuse: LangfuseConfigandopenTelemetry: OpenTelemetryConfigfields are removed. A backend reads its block fromextras-LangfuseConfig.fromExtras(settings.extras),OpenTelemetryConfig.fromExtras(settings.extras)- and code that built settings by hand passesextras = Map("publicKey" -> ..., "secretKey" -> ...)for Langfuse, orMap("serviceName" -> ..., "endpoint" -> ..., "headers.Authorization" -> ...)for OpenTelemetry.LangfuseConfigmoves tollm4s-observability, andOpenTelemetryConfigtollm4s-observability-otel, in the same package.DefaultConfigis removed. Its last four constants moved toLangfuseConfig:DEFAULT_LANGFUSE_URL,_ENV,_RELEASEand_VERSIONareLangfuseConfig.DEFAULT_URL,DEFAULT_ENV,DEFAULT_RELEASEandDEFAULT_VERSION.ConfigKeys.LANGFUSE_*areLangfuseConfigKeys.LANGFUSE_*inllm4s-observability.RAGASLangfuseObserver.fromTracingSettings(TracingSettings)is removed: it read the removedTracingSettings.langfuse. UseRAGASLangfuseObserver.from(config)withLangfuseConfigLoader.default(), as above.TraceEvent.createTraceEventandorg.llm4s.llmconnect.model.TraceHelperare removed. Both built Langfuse ingestion JSON - a"trace-create"batch envelope, andevent-create/generation-create/span-createenvelopes per conversation message - and nothing in llm4s called either;LangfuseTracingbuilds its own batches. They were Langfuse’s wire format sitting in the core contract, so they are deleted rather than moved. Code that called them builds the JSON itself, or traces throughLangfuseTracing(llm4s-observability), which sends the same event types.
Slice 6: tracing backends are discovered, and agent state is a TraceEvent
The first slice 6 change (#1133) lands the extension
point for tracing before Langfuse and Prometheus are carved out of llm4s-core, as slice 4 did for
providers. Nothing moves module yet, and nothing in a release has it; 0.4.1 and earlier behave as
before.
Tracing.traceAgentState(AgentState) is removed
Tracing no longer mentions AgentState, so the tracing contract does not depend on the agent
runtime. Agent state is traced as an ordinary event, TraceEvent.AgentStateUpdated, which
AgentState#toTraceEvent builds:
1
2
3
4
5
// before
tracing.traceAgentState(state)
// after
tracing.traceEvent(state.toTraceEvent)
AgentStateUpdated gained a messages: Seq[Message] field (default empty) before timestamp,
carrying the conversation, so that Langfuse still records a trace with one span per message. It is
not part of toJson. The agent emits the event exactly where it called traceAgentState.
A custom Tracing implementation drops its traceAgentState override; anything it did there
belongs in traceEvent’s AgentStateUpdated case.
Tracing.create finds backends on the classpath
Tracing.create(settings) keeps its signature. It builds NoOp and Console itself and hands
every other mode to an org.llm4s.trace.spi.TracingBackend registered in
META-INF/services/org.llm4s.trace.spi.TracingBackend:
- OpenTelemetry:
llm4s-observability-otelregistersOpenTelemetryTracingBackend. TheClass.forNamereflection core used to find it is gone. Adding the dependency is all it takes, as before. - Langfuse: registered by core’s own services entry until it is carved into
llm4s-observability- which has since happened; see the carve’s note. - Anything else: implement
TracingBackend(aclasswith a public no-arg constructor, not anobject) withmode = TracingMode.Named("yourmode"), declare it in the services file, andTRACING_MODE=yourmodeselects it.
Tracing.create still never fails: a missing or broken backend logs an error and gives
NoOpTracing. The new Tracing.fromSettings(settings): Result[Tracing] returns that as a
ConfigurationError (or the backend’s own error) instead, and
Tracing.fromSettings(settings, TracingBackends.of(...)) registers a backend explicitly, for a
shaded jar whose services files did not survive.
TracingMode is open
TracingMode gained a Named(name) case and a name member. TracingMode.fromString returns
Named for a value it does not recognise, lower-cased, where it used to return NoOp with a
warning; a blank value is still NoOp. The end result of an unknown TRACING_MODE is unchanged -
Tracing.create gives NoOpTracing - but the log line is now an error that lists the modes that
are available.
Behaviour changes
- An
AgentStateUpdatedwith no messages is exported to Langfuse as a summary trace. The oldtraceAgentStatesent nothing for an empty conversation. - Langfuse reports a failed export of the conversation trace as a
Left;traceAgentStatealways returnedRight(()). The agent swallows tracing errors either way. -
The agent state span takes the event’s name. Agent runs used to reach a tracer through
traceAgentState, whose span name differed from the onetraceEventgave the same event; they now go throughtraceEvent, so the attributes are unchanged but the name is not. Update any dashboard or query keyed on the old name:Backend Old name (agent runs) New name OpenTelemetry Agent State SnapshotAgent State UpdatedTraceCollectorTracingagent-state-updateagent_state_updated - An OpenTelemetry SDK that fails to start is reported by
Tracing.fromSettingsand givesNoOpTracingfromTracing.create, rather than a tracer whose every call failed.
Source breaks
Tracing.traceAgentStateis removed - usetraceEvent(state.toTraceEvent).TraceEvent.AgentStateUpdatedhas a fifth field,messages, beforetimestamp. A positionalAgentStateUpdated(status, messages, logs, timestamp)must name the timestamp (timestamp = ...), and a patternAgentStateUpdated(a, b, c, ts)needs one more binder.TracingModehas a new case,Named, so an exhaustivematchover it needs one more branch, and a new abstract member,name, for anything extending the sealed trait (nothing outside core can).TracingMode.fromStringreturnsNamed(...), notNoOp, for an unrecognised value.
From LLM_MODEL to named provider sections
API keys: superseded by Vendor credentials. Provider modules now bind each vendor’s API-key variable (
OPENAI_API_KEY, …), so theapiKey = ${?OPENAI_API_KEY}lines below, and the embeddings binding, are optional. What this note says aboutLLM_MODELand base-URL variables still holds.
Since #903 (in 0.3.2) removed legacy single-provider
loading, nothing in llm4s reads LLM_MODEL, nor a provider’s API-key or base-URL variable
(OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY, AZURE_API_BASE, OLLAMA_BASE_URL, …).
No reference.conf binds them. Setting them has no effect, and an application that relied on them
fails at startup with a configuration error. The documentation kept teaching them until
#1132’s follow-up.
Configure providers as named sections in your own application.conf, binding secrets from the
environment with ${?VAR} - you choose the variable names:
1
2
3
4
5
6
7
8
9
10
11
12
# src/main/resources/application.conf
llm4s {
providers {
provider = "openai-main" # the default: the name of a section below
openai-main {
provider = "openai"
model = "gpt-4o-mini"
apiKey = ${?OPENAI_API_KEY}
}
}
}
| Before | After |
|---|---|
LLM_MODEL=openai/gpt-4o-mini |
a section with provider = "openai", model = "gpt-4o-mini", selected by llm4s.providers.provider |
OPENAI_API_KEY=... read automatically |
apiKey = ${?OPENAI_API_KEY} in that section, and the same variable exported |
OPENAI_BASE_URL, OLLAMA_BASE_URL, … |
baseUrl = "..." (or baseUrl = ${?OLLAMA_BASE_URL}) in the section; Ollama sections require it |
switching model by changing LLM_MODEL |
a second section, selected by changing provider, by -Dllm4s.providers.provider=<name>, by your own provider = ${?LLM4S_PROVIDER} binding, or loaded directly with Llm4sConfig.provider("<name>") |
Llm4sConfig.provider() |
Llm4sConfig.defaultProvider() |
Two things to know:
- Only the section you load is validated (since
#1132’s follow-up; not in a release yet). A
section whose required
apiKeyis unset, or whose provider module is not on the classpath, failsprovider("<that section>")- anddefaultProvider()when it is the default - but not a load of any other section. Up to 0.4.1 every section was validated on every load, so each environment had to fill in every section or ship a config without it.Llm4sConfig.providers(), which returns every section, still validates them all. - OpenAI embeddings do not see
OPENAI_API_KEYeither. WithEMBEDDING_MODEL=openai/<model>(which is bound, by llm4s-core’sreference.conf), addllm4s.embeddings.openai.apiKey = ${?OPENAI_API_KEY}. Since #1132’s follow-up the key no longer falls back tollm4s.openai.apiKey, the pre-#903 chat key that nothing else read: if you set that, setllm4s.embeddings.openai.apiKeyinstead.
Variables that reference.conf files do bind - TRACING_MODE, LANGFUSE_*, OTEL_SERVICE_NAME,
OTEL_EXPORTER_OTLP_ENDPOINT, EMBEDDING_MODEL, VOYAGE_API_KEY and others - still work; see
the variables llm4s reads.
Two tools read LLM_MODEL themselves: the chat-tui sample (ChatTuiConfig) and the config-policy
env check (EnvCheckPolicies).
Named providers: provider-specific keys; Vertex AI project and location
A named provider section can now carry keys of the provider’s own, declared by the provider rather than squeezed into the shared fields (#1215). Not in a release yet.
Vertex AI: endpoint becomes project, organization becomes location
Vertex AI used to read its GCP project from endpoint and its region from organization. It
now has keys named for what they are:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
# before
vertex-main {
provider = "vertexai"
model = "gemini-2.0-flash"
endpoint = "my-gcp-project"
organization = "europe-west4"
}
# after
vertex-main {
provider = "vertexai"
model = "gemini-2.0-flash"
project = "my-gcp-project"
location = "europe-west4" # optional, default us-central1
}
The old spellings still work for now, as deprecated aliases: each one logs a warning naming the
section and the key to rename it to, and will stop working in a later release. Setting both
project and a different endpoint (or location and a different organization) is an
error rather than a guess. A missing project is now reported as
- project: the GCP project ID that owns your Vertex AI resources (set it in application.conf under llm4s.providers.<name>.project).
Unknown keys are reported
A key in a section that is neither a built-in field nor one the provider declares used to be
ignored silently. It is still ignored - existing configs keep loading - but now with a warning
that names it and lists the keys the provider accepts, which is how a typo like regoin shows
up. An extra key whose value is an object or list is an error, since none takes one.
The missing-baseUrl hint no longer suggests a variable nothing reads
A section missing a required baseUrl used to say set <PROVIDER>_BASE_URL, for example
set OLLAMA_BASE_URL, though nothing read such a variable for a named section. It now says
where to set the key - baseUrl: set it in application.conf under llm4s.providers.<name>.baseUrl (e.g. ...).
Where the provider has a conventional variable (ProviderConfigSpec.baseUrlEnv), the message
shows the binding that makes it read, since a named section reads none by itself:
openai-compatible adds to read it from OPENAI_COMPATIBLE_BASE_URL, add baseUrl = ${?OPENAI_COMPATIBLE_BASE_URL} to the section.
Code matching on the old message text needs updating.
Source breaks
RawNamedProviderSectionandNamedProviderConfiggained a trailingextrasparameter with a default, so construction by name or position still compiles; pattern matches that list every field (case NamedProviderConfig(a, b, ...)) need one more.ProviderConfigSpecgainedbaseUrlEnvandextras, both defaulted.VertexAIProvider.configSpecno longer setsrequiresEndpoint. Code that builds aNamedProviderConfigby hand and passes it straight toVertexAIProvider.buildConfig, skipping validation, must setextras = Map("project" -> ..., "location" -> ...): the deprecated aliases are resolved by validation, not bybuildConfig.
For provider authors
Declare provider-specific keys in ProviderConfigSpec.extras and read them from
NamedProviderConfig.extras - see CONTRIBUTING.
Slice 5 follow-ups: llm4s-openai-compatible
Fixes to the shared OpenAI-compatible client after its carve (#1132). Nothing here needs a code change unless you depend on the old behaviour.
complete times out after two minutes
OpenAICompatibleClient.complete - DeepSeek, Z.ai, OpenRouter, Mistral, Cohere and the generic
openai-compatible provider - sent its request with no timeout, so an endpoint that never
answered hung the caller. It now fails after two minutes (OpenAICompatibleClient.RequestTimeout),
the value the old Mistral and Cohere clients used; streaming keeps five minutes. Neither is
configurable yet (#712): a slow local model with a
long prompt should stream.
Streaming requests send stream_options
A streaming request from the generic provider and DeepSeek now carries
"stream_options": {"include_usage": true}, so servers that follow OpenAI - vLLM, Ollama’s
/v1, Perplexity’s Router - report token usage on streams. Z.ai, OpenRouter, Mistral and Cohere
do not send it. An endpoint that rejects unknown fields fails the stream with a 400 or 422; build
its config with OpenAICompatibleConfig.fromValues(..., streamUsage = false). A dialect you wrote
yourself inherits streamUsageOption = true; override it to false if the provider rejects the
field.
Requesty configs report requesty
A config loaded from a provider = "requesty" section reported providerId = openai, because
OpenAIConfig inferred its id from the base URL. OpenAIConfig now carries the id its descriptor
sets, in a new trailing field explicitProviderId: Option[ProviderId] = None
(OpenAIConfig.fromValues has a matching defaulted providerId parameter), so it reports
requesty; provider = "openrouter" sets openrouter the same way, so an OpenRouter section
with a proxy baseUrl still reaches OpenRouter. With None - a config built by hand, or
provider = "openai" - the id is inferred from the base URL as before.
What you may notice:
llm4s-config-policyseesrequesty: anallowedProviderslist must namerequestyrather than rely onopenai, and model patterns matchrequesty/<model>.LLMConnect.getClient(config)on a Requesty config now dispatches to therequestydescriptor, which needsllm4s-openai- as theopenaione did.OpenAIConfig.toStringshowsproviderId; a pattern match destructuring all six fields must add a seventh.
OPENAI_COMPATIBLE_BASE_URL and OPENAI_COMPATIBLE_API_KEY
These are now the conventional variables for a generic endpoint, named by
OpenAICompatibleConfigKeys. Llm4sConfig does not read them - the generic provider has no vendor
and so no shared credentials key, and nothing reads LLM_MODEL - so bind them in a section (baseUrl = ${?OPENAI_COMPATIBLE_BASE_URL}). The
chat-tui sample accepts LLM_MODEL=openai-compatible/<model> with them, and the config-policy
env check reads OPENAI_COMPATIBLE_BASE_URL as that provider’s endpoint instead of
OPENAI_BASE_URL.
Slice 5: Mistral, Cohere and Voyage leave core; core ships no provider
The last provider clients leave llm4s-core (#1132).
None of this is in a release yet; 0.4.1 and earlier still ship Mistral, Cohere and Voyage inside
llm4s-core.
| Provider | Now in | Add |
|---|---|---|
Mistral (provider = "mistral") |
llm4s-openai-compatible, as a dialect |
"org.llm4s" %% "llm4s-openai-compatible" |
Cohere (provider = "cohere") |
llm4s-openai-compatible, as a dialect |
"org.llm4s" %% "llm4s-openai-compatible" |
Voyage AI embeddings (EMBEDDING_MODEL=voyage/...) |
llm4s-voyage (modules/providers/voyage) |
"org.llm4s" %% "llm4s-voyage" |
Package names are unchanged, and so are the names and constructors of MistralClient,
MistralProvider, MistralConfig, CohereClient, CohereProvider, CohereConfig and
VoyageAIEmbeddingProvider, so imports compile as before once the dependency is added. No
environment variable or config key changed its name; llm4s.embeddings.voyage moved to
llm4s-voyage’s reference.conf, so it exists exactly when the module is on the classpath.
Mistral and Cohere stream now
Both are OpenAI-compatible, so they are small dialects on
the shared OpenAICompatibleClient rather than clients of their own. Both gain streaming
(streamComplete returned “not supported” - #925),
streamed tool calls and token usage, tool calling and structured output; their descriptors no
longer declare streaming = false.
- Mistral posts to
<baseUrl>/v1/chat/completionsas before (baseUrlis the API root,https://api.mistral.ai); abaseUrlalready ending in/v1is no longer doubled. Tool-call ids are sent in the nine-character form Mistral insists on; a reasoning model’s thinking is returned asCompletion.thinking. - Cohere now calls Cohere’s
OpenAI-compatibility API instead of the native
/v2/chat.CohereConfig.DEFAULT_BASE_URLis nowhttps://api.cohere.ai/compatibility/v1(it washttps://api.cohere.com). A configuredbaseUrlkeeps working: one that does not end in/compatibility/v1is taken to be a Cohere API root, as it was, and gets/compatibility/v1appended (after a trailing/v1or/v2is dropped);CohereConfig.fromValuesstores and logs the mapped URL, soendpointUrland policy checks see where requests really go. A proxy that forwarded only/v2/chatmust now forward/compatibility/v1/chat/completions. System messages go under thedeveloperrole and a JSON schema as{"type": "json_object", "schema": ...}, as Cohere documents.
Behaviour changes for both, now that they share the client DeepSeek, Z.ai and OpenRouter use:
- Reply text is not trimmed (the old clients trimmed it).
- A reply with no text is an empty completion, not a
ValidationError- with tool calling, a reply may carry only tool calls. - A missing
idorcreatedis left""/0rather than a random UUID / the current time (Mistral), and Cohere’screatednow comes from the reply. - A
ToolMessageis sent rather than refused (Mistral) or silently dropped (Cohere). CompletionOptions.reasoningis still not sent to either: Mistral’sreasoning_effortand Cohere’s accept only some models or values.
The shared client itself changed for every provider on it: streamed completions now report
token usage and a cost estimate (they always came back with usage = None), and an empty
conversation fails with a ValidationError before any request is sent.
Core ships no provider
With those three gone, llm4s-core holds no provider client, and the list that existed only
because it did is removed:
BuiltinProvidersandBuiltinProviderModule(org.llm4s.llmconnect.provider) are deleted, with core’sMETA-INF/services/org.llm4s.llmconnect.spi.Llm4sProviderModuleentry.-
ProviderRegistry.builtinis deleted. It would now be empty. Where discovery cannot run - a fat jar whose services files were overwritten - name the provider modules you ship:1 2
given ProviderRegistry = ProviderRegistry.ofModules(new Llm4sOpenAIModule, new Llm4sOpenAICompatibleModule, new Llm4sVoyageModule)
ProviderRegistry.default (discovery) is unchanged, and so is every call site that relies on it.
Core’s llmconnect/provider package now holds only provider-neutral helpers: CostEstimator,
EmbeddingProvider, HttpErrorMapper, MetricsRecording, ProviderExchangeRecorder and
ProviderResultOps.
Source breaks
BuiltinProviders,BuiltinProviderModuleandProviderRegistry.builtinare removed - see above.ProviderModelListers.Mistralis nowMistralModelLister(org.llm4s.config, inllm4s-openai-compatible);MistralProvider.modelListerreturns it.ConfigKeys.MISTRAL_API_KEYandMISTRAL_BASE_URLare now onOpenAICompatibleConfigKeys, andConfigKeys.VOYAGE_API_KEY,VOYAGE_EMBEDDING_BASE_URLandVOYAGE_EMBEDDING_MODELonVoyageConfigKeys(llm4s-voyage), both inorg.llm4s.config. The strings are unchanged.CohereConfig.DEFAULT_BASE_URLchanged value - see above.OpenAICompatibleDialectgained four members -sendEmptyAssistantTurns,encodeToolCallId,systemRoleandencodeResponseFormat- each defaulting to the standard format, so an existing dialect compiles and behaves as before.
Slice 5: llm4s-openai-compatible
The fifth provider module of slice 5 (#1132),
carrying DeepSeek (provider = "deepseek"), Z.ai ("zai"), OpenRouter ("openrouter") and a
new generic provider, "openai-compatible", for any other endpoint that speaks the OpenAI
/chat/completions API. It is in the build but not yet in a release; 0.4.1 and earlier still
ship DeepSeek, Z.ai and OpenRouter inside llm4s-core, and have no generic provider.
Unlike the earlier carves this is a consolidation, not a pure move. DeepSeek, Z.ai and
OpenRouter each had their own ~400-line copy of the same SDK-free client; they are now thin
subclasses of one OpenAICompatibleClient, each with a small OpenAICompatibleDialect for what
genuinely differs (headers, content encoding, reasoning parameters, thinking extraction,
tool-call parsing). The module depends on nothing but llm4s-core.
1
libraryDependencies += "org.llm4s" %% "llm4s-openai-compatible" % version
What moved
| Code | Now in |
|---|---|
DeepSeekClient, DeepSeekProvider, ZaiClient, ZaiProvider, OpenRouterClient, OpenRouterProvider (org.llm4s.llmconnect.provider) |
llm4s-openai-compatible |
DeepSeekConfig, ZaiConfig, OpenAIConfig (org.llm4s.llmconnect.config) |
llm4s-openai-compatible |
ProviderModelListers.DeepSeek / .OpenRouter → DeepSeekModelLister / OpenRouterModelLister (org.llm4s.config) |
llm4s-openai-compatible |
DefaultConfig.DEFAULT_DEEPSEEK_BASE_URL → DeepSeekConfig.DEFAULT_BASE_URL |
llm4s-openai-compatible |
DefaultConfig.DEFAULT_OPENROUTER_BASE_URL → OpenRouterProvider.DEFAULT_BASE_URL |
llm4s-openai-compatible |
ConfigKeys.DEEPSEEK_API_KEY, DEEPSEEK_BASE_URL, OPENROUTER_BASE_URL → OpenAICompatibleConfigKeys (org.llm4s.config) |
llm4s-openai-compatible |
the commented deepseek-main, zai-main and openrouter-main examples in reference.conf |
llm4s-openai-compatible’s reference.conf |
New: OpenAICompatibleClient, OpenAICompatibleDialect, OpenAICompatibleConfig,
OpenAICompatibleProvider, OpenAICompatibleModelLister, Llm4sOpenAICompatibleModule.
Package names are unchanged, and so are the public shapes of the three clients - their
constructors and companion apply overloads - their descriptors and their configs, so
new DeepSeekClient(config) or OpenRouterClient(config, metrics) compile as before once the
dependency is added.
llm4s-openai users: OpenAIConfig moved here, and llm4s-openai now depends on
llm4s-openai-compatible to get it. That module brings no SDK, and nothing changes in your build.
The generic openai-compatible provider
1
2
3
4
5
6
7
8
9
10
11
llm4s.providers {
local-vllm {
provider = "openai-compatible"
baseUrl = "http://localhost:8000/v1" # required
model = "Qwen/Qwen2.5-7B-Instruct" # required
# apiKey = ... # optional; no Authorization header without one
# contextWindow = 32768 # optional; default 8192
# reserveCompletion = 4096 # optional; default 2048
# headers { X-Team = "search" } # optional
}
}
To support it, a named provider section may now carry contextWindow, reserveCompletion and a
headers object, which NamedProviderConfig exposes. Providers other than openai-compatible
ignore them. See OpenAI-compatible endpoints.
Registration is the dependency
llm4s-openai-compatible declares Llm4sOpenAICompatibleModule in its META-INF/services, so
ProviderRegistry.default finds it. Without the dependency, provider = "deepseek", "zai" and
"openrouter" fail with the registry’s error, which names the providers that are registered.
ProviderRegistry.builtin no longer includes them; where discovery cannot run:
1
given ProviderRegistry = ProviderRegistry.ofModules(new Llm4sOpenAICompatibleModule) // `builtin` was removed later in slice 5
Behaviour changes
The three copies had drifted apart; the shared client does each thing one way:
- DeepSeek returns thinking.
deepseek-reasoner’sreasoning_contentis nowCompletion.thinkingand is streamed as thinking deltas; the old client dropped it. Itscompletion_tokens_details.reasoning_tokensisTokenUsage.thinkingTokens. - The stream body is closed on every failure. Z.ai and OpenRouter left it open on an error status.
- Every call records exactly one provider exchange, including a request that cannot be sent (DeepSeek and Z.ai recorded none for a non-streaming one; OpenRouter recorded failures twice).
- OpenRouter sends assistant content as a string. The old client passed an
Optionthrough ujson’s implicit conversion, so"hi"went out as["hi"]and no content as[]; it is now"hi",""ornull. - Z.ai reads usage given as an array, which its client meant to support but never matched.
- Streamed tool calls keep all their arguments. A tool call streamed across several deltas
lost every fragment after the first in all three clients: continuations carry only an
index, the missing id was defaulted to"", andStreamingAccumulatorskips a chunk with no id. The shared client now maps each index to its call’s id for the life of the stream, and a streamedCompletionreports its tool calls intoolCalls, as a non-streaming one does. - A reply’s
message.contentOptisNonewhen the reply has no text for all three (it wasSome("")for DeepSeek and Z.ai);Completion.contentis""either way. - Replies are read leniently where the copies threw: a missing
id,createdormodeldefaults, a streamed event with nochoicesis skipped, and a malformed non-streaming reply is aLeftfor all three (Z.ai could throw). OpenRouter keeps its strict tool-call parsing; DeepSeek and Z.ai keep their lenient one.
Source breaks
ProviderModelListers.DeepSeekand.OpenRouterare nowDeepSeekModelListerandOpenRouterModelLister, in the same package. The descriptors’modelListerreturns them.DefaultConfig.DEFAULT_DEEPSEEK_BASE_URLandDEFAULT_OPENROUTER_BASE_URLare nowDeepSeekConfig.DEFAULT_BASE_URLandOpenRouterProvider.DEFAULT_BASE_URL. The values are unchanged.ConfigKeys.DEEPSEEK_API_KEY,DEEPSEEK_BASE_URLandOPENROUTER_BASE_URLare now onOpenAICompatibleConfigKeys, in the same package. The strings are unchanged.ProviderRegistry.builtinno longer includesdeepseek,zaioropenrouter- see above.NamedProviderConfigandRawNamedProviderSectiongained three trailing fields with defaults (contextWindow,reserveCompletion,headers), so construction by name or by position is unaffected; a pattern match that destructures all seven fields must add three.NamedProviderConfig.toStringnow redacts the API key and header values.-
ProviderModelListers.openAICompatiblegained two defaulted parameters,extraHeadersandapiKeyRequired; existing calls compile unchanged. It no longer special-cases OpenRouter, whose lister passes its headers explicitly. OpenRouterToolCallDeserializeris removed fromorg.llm4s.llmconnect.serialization. No client used it after the consolidation. The “double-nested” array it parsed was an artefact of the oldOpenRouterClient, not OpenRouter’s format; useStandardToolCallDeserializer(which stays), asOpenRouterClientnow does.StreamingResponseHandleris removed, withBaseStreamingResponseHandler,OpenAIStreamingHandler,AnthropicStreamingHandlerandStreamingResponseHandler.forProvider(org.llm4s.llmconnect.streaming). No client streamed through them - each parses its own stream and accumulates withStreamingAccumulator, which stays - andforProviderwas called only from tests. To assemble streamed chunks yourself, feed them to aStreamingAccumulator.
What did not change
Every configuration key and environment variable for DeepSeek, Z.ai and OpenRouter: their
provider ids, apiKey, baseUrl and organization, and their default base URLs.
Slice 5: llm4s-openai moves to openai-java
llm4s-openai’s OpenAIClient - behind provider = "openai", "azure" and "requesty" - now
runs on OpenAI’s official Java SDK, com.openai:openai-java, instead of Microsoft’s
com.azure:azure-ai-openai. Microsoft has
deprecated that SDK
(its last release, 1.0.0-beta.16, was on 2025-03-26) and points to openai-java, which also
covers Azure OpenAI (#1132).
Most users change nothing. OpenAIClient’s constructors and apply overloads,
OpenAIProvider, AzureProvider, RequestyProvider, OpenAIConfig and AzureConfig, every
configuration key and every environment variable are unchanged. Your dependency tree changes:
llm4s-openai now brings OkHttp, Jackson (2.x) and the Kotlin standard library rather than the
Azure core libraries. openai-java checks its Jackson version when a client is built and fails
on a Jackson it cannot use (a different major version, anything before 2.13.4, or 2.18.1); if
your application pins Jackson, keep it on a compatible 2.x.
How Azure is configured
As before: endpoint is the resource endpoint (https://<resource>.openai.azure.com), model
is the deployment name, apiKey is sent as the api-key header, and apiVersion as the
api-version query parameter, so a request goes to
<endpoint>/openai/deployments/<deployment>/chat/completions?api-version=<version>. The client
tells the SDK this is Azure rather than letting it guess from the host name, so an endpoint on
your own domain (API Management, a private endpoint) keeps working. Two things are new:
apiVersiontakes either form. The wire form (2024-10-21,2025-01-01-preview), which the docs have always shown, now works; before, only the Azure SDK’s constant names (V2024_10_21,V2025_01_01_PREVIEW- the form ofAzureConfig.DEFAULT_API_VERSION) did. Both still work.- An endpoint ending in
/openai/v1uses Azure’s unified v1 API: the deployment goes in the request body andapi-versionis sent only if you set one other than the default.
Source break: AzureToolHelper is now OpenAIToolHelper
AzureToolHelper took and returned Azure SDK types, so it could not survive the SDK. It is
replaced, in the same package (org.llm4s.toolapi) and module, by OpenAIToolHelper over
openai-java’s types (com.openai.models.chat.completions):
Before (AzureToolHelper) |
Now (OpenAIToolHelper) |
|---|---|
addToolsToOptions(registry, options: ChatCompletionsOptions): ChatCompletionsOptions |
addToolsToParams(registry, builder: ChatCompletionCreateParams.Builder): ChatCompletionCreateParams.Builder |
convertToolRegistryToAzureTools(registry): java.util.List[ChatCompletionsToolDefinition] |
convertToolRegistryToOpenAITools(registry): java.util.List[ChatCompletionTool] |
1
2
3
4
5
import org.llm4s.toolapi.{ OpenAIToolHelper, ToolRegistry }
import com.openai.models.chat.completions.ChatCompletionCreateParams
val params = OpenAIToolHelper
.addToolsToParams(new ToolRegistry(tools), ChatCompletionCreateParams.builder().model("gpt-4o"))
If you called the Azure SDK yourself alongside llm4s, add com.azure:azure-ai-openai to your
own build: llm4s-openai no longer brings it.
Behaviour changes
- Streamed tool calls keep their arguments. A streamed tool call arrives split across
deltas, and only the first carries the call’s
id; the Azure SDK path keyed calls byid, so every later fragment - usually all of the arguments - was lost. Continuations are now matched byindex, fragments are concatenated verbatim, and a streamedCompletionreports its tool calls intoolCallsas a non-streaming one does. OpenAIConfig.organizationis sent as theOpenAI-Organizationheader. The Azure SDK ignored it.- HTTP errors map by status code: 401 and 403 to
AuthenticationError, 429 toRateLimitError, 400 toValidationError, anything else toServiceError, each naming the provider. Before, they were classified by searching the exception message, and most becameUnknownError. - Streamed token usage is read from whichever chunk carries it, including a usage-only final chunk, not only from the chunk with the finish reason.
close()releases the SDK’s HTTP client (connections and threads); before it released nothing.- Azure and Requesty are labelled as themselves. Their errors (
AuthenticationError.providerand the error context), metrics and provider-exchange log now sayazureandrequesty; every one saidopenaibefore. Requesty takes its label from its descriptor. A Requesty config loaded from aprovider = "requesty"section also reportsproviderId=requestysince a later fix (see above); only anOpenAIConfigyou build by hand with Requesty’s base URL still infersopenaifrom it - passproviderId = Some(ProviderId("requesty"))toOpenAIConfig.fromValuesfor that. - Several streamed tool calls come back in the order the stream named them, in
Completion.toolCallsand on the message.llm4s-core’sStreamingAccumulatorkept them in an unordered map, so they could come back in hash order; this applies to every client that streams through it.
Slice 5: llm4s-openai
The fourth provider module of slice 5 (#1132),
carrying the three providers that share OpenAIClient - OpenAI (provider = "openai"), Azure
OpenAI ("azure") and Requesty ("requesty") - and the OpenAI embedding provider
(EMBEDDING_MODEL=openai/<model>). It is in the build but not yet in a release; 0.4.1 and
earlier still ship them inside llm4s-core.
With it goes the Azure OpenAI SDK (com.azure:azure-ai-openai), which OpenAIClient is built
on: llm4s-core no longer depends on it, and with the Anthropic SDK already gone, core now
depends on no vendor SDK at all.
OpenRouter, DeepSeek and Z.ai are not in this module. They speak the OpenAI wire format but
each has its own client with no SDK, so bundling them here would make their users download the
Azure SDK for nothing. They went on to llm4s-openai-compatible - see
above.
What moved
| Code | Now in |
|---|---|
OpenAIClient, OpenAIProvider, AzureProvider, RequestyProvider, OpenAIEmbeddingProvider (org.llm4s.llmconnect.provider) |
llm4s-openai |
AzureConfig (org.llm4s.llmconnect.config) |
llm4s-openai |
AzureToolHelper (org.llm4s.toolapi) |
llm4s-openai |
ProviderModelListers.OpenAI / .Requesty → OpenAIModelLister / RequestyModelLister (org.llm4s.config) |
llm4s-openai |
DefaultConfig.DEFAULT_OPENAI_BASE_URL → OpenAIProvider.DEFAULT_BASE_URL |
llm4s-openai |
DefaultConfig.DEFAULT_REQUESTY_BASE_URL → RequestyProvider.DEFAULT_BASE_URL |
llm4s-openai |
DefaultConfig.DEFAULT_AZURE_V2025_01_01_PREVIEW → AzureConfig.DEFAULT_API_VERSION |
llm4s-openai |
ConfigKeys.OPENAI_*, REQUESTY_BASE_URL, AZURE_*, OPENAI_EMBEDDING_* → OpenAIConfigKeys (org.llm4s.config) |
llm4s-openai |
the commented openai-main, requesty-main and azure-main examples, and the llm4s.embeddings.openai block, in reference.conf |
llm4s-openai’s reference.conf |
Package names are unchanged, so import org.llm4s.llmconnect.provider.OpenAIClient keeps
working once the dependency is added:
1
libraryDependencies += "org.llm4s" %% "llm4s-openai" % version
What stayed in core
OpenAIConfig, because OpenRouter builds one too:OpenRouterProviderandOpenRouterClienttake anOpenAIConfig, and itsproviderIdanswersopenrouterfor an OpenRouter base URL. It moved when OpenRouter did, tollm4s-openai-compatible, whichllm4s-openainow depends on.OpenAIStreamingHandler(since removed withStreamingResponseHandler; seellm4s-openai-compatible), the SSE parser behindStreamingResponseHandler.forProvider("openai" | "azure" | "openrouter"), which OpenRouter’s path shares.OpenAIClientstreams through the Azure SDK.ConfigKeys.OPENROUTER_BASE_URL, still namingOPENAI_BASE_URL(since moved toOpenAICompatibleConfigKeys).- Strings that do not reach a client:
ToolRegistry.getOpenAIToolsandgetToolDefinitionsSafe("openai"), theopenai/...model-registry data, thesk-secret pattern, config-policy allow-lists.
Registration is the dependency
llm4s-openai declares Llm4sOpenAIModule in its META-INF/services, so
ProviderRegistry.default finds it and provider = "openai", "azure" and "requesty", and
EMBEDDING_MODEL=openai/<model>, resolve as before. Without the dependency they fail with the
registry’s error, which says the provider is not registered and names the providers that are.
ProviderRegistry.builtin no longer includes them. If you used builtin to avoid classpath
discovery (a shaded fat jar, typically), add the module explicitly:
1
given ProviderRegistry = ProviderRegistry.ofModules(new Llm4sOpenAIModule) // `builtin` was removed later in slice 5
llm4s-rag users: RAGConfig.default embeds with openai/text-embedding-3-small. A
pipeline built from the default therefore needs llm4s-openai too; otherwise name the provider
you ship with .withEmbeddings("voyage", ...) (or "ollama" with llm4s-ollama).
llm4s-rag does not depend on llm4s-openai itself, so it does not bring the Azure SDK.
Source breaks
ToolRegistry.addToAzureOptions(options)is removed. Its signature exposed the Azure SDK typeChatCompletionsOptionsfrom core’sToolRegistry, so core could not drop the SDK while it existed. CallAzureToolHelper.addToolsToOptions(registry, options)instead - same package (org.llm4s.toolapi), now inllm4s-openai, with the same result.ProviderModelListers.OpenAIand.Requestyare nowOpenAIModelListerandRequestyModelLister, in the same package (org.llm4s.config). The descriptors’modelListerreturns them, so code that reached a lister through the descriptor is unaffected.- Three defaults moved off
DefaultConfig:DEFAULT_OPENAI_BASE_URLis nowOpenAIProvider.DEFAULT_BASE_URL,DEFAULT_REQUESTY_BASE_URLis nowRequestyProvider.DEFAULT_BASE_URL, andDEFAULT_AZURE_V2025_01_01_PREVIEWis nowAzureConfig.DEFAULT_API_VERSION. The values are unchanged. The base URLs live on the descriptors rather than onOpenAIConfigbecauseOpenAIConfigstays in core. ConfigKeys.OPENAI_API_KEY,OPENAI_BASE_URL,OPENAI_ORG,REQUESTY_BASE_URL,AZURE_API_BASE,AZURE_API_KEY,AZURE_API_VERSION,OPENAI_EMBEDDING_BASE_URLandOPENAI_EMBEDDING_MODELare now onOpenAIConfigKeys, in the same package, asConfigKeys.ANTHROPIC_*becameAnthropicConfigKeys. The strings are unchanged.ProviderRegistry.builtinno longer includesopenai,azureorrequesty, nor theopenaiembedding provider - see above.
What did not change
Every configuration key and environment variable: llm4s.providers.<name> with
provider = "openai", "azure" or "requesty" and their apiKey, baseUrl, organization,
endpoint and apiVersion; llm4s.embeddings.openai.*, OPENAI_EMBEDDING_BASE_URL and
OPENAI_EMBEDDING_MODEL; and llm4s.openai.apiKey as the key OpenAI embeddings share with chat.
(A later change dropped that llm4s.openai.apiKey fallback: the key is
llm4s.embeddings.openai.apiKey alone - see
From LLM_MODEL to named provider sections.)
Slice 5: llm4s-anthropic
The third provider module of slice 5 (#1132),
carrying the Anthropic Claude chat provider (provider = "anthropic"). It is in the build but
not yet in a release; 0.4.1 and earlier still ship it inside llm4s-core.
With it goes the Anthropic Java SDK (com.anthropic:anthropic-java): llm4s-core no longer
depends on it, so an application that does not use Anthropic no longer carries it. It was also
declared, unused, by llm4s-workspace-client, and has been removed from there too.
What moved
| Code | Now in |
|---|---|
AnthropicClient, AnthropicProvider (org.llm4s.llmconnect.provider) |
llm4s-anthropic |
AnthropicConfig (org.llm4s.llmconnect.config) |
llm4s-anthropic |
ProviderModelListers.Anthropic → AnthropicModelLister (org.llm4s.config) |
llm4s-anthropic |
DefaultConfig.DEFAULT_ANTHROPIC_BASE_URL → AnthropicConfig.DEFAULT_BASE_URL |
llm4s-anthropic |
ConfigKeys.ANTHROPIC_API_KEY, ConfigKeys.ANTHROPIC_BASE_URL → AnthropicConfigKeys (org.llm4s.config) |
llm4s-anthropic |
the commented anthropic-main example in reference.conf |
llm4s-anthropic’s reference.conf |
Package names are unchanged, so import org.llm4s.llmconnect.config.AnthropicConfig keeps
working once the dependency is added:
1
libraryDependencies += "org.llm4s" %% "llm4s-anthropic" % version
Registration is the dependency
llm4s-anthropic declares Llm4sAnthropicModule in its META-INF/services, so
ProviderRegistry.default finds it and provider = "anthropic" resolves as before. Without the
dependency it fails with the registry’s error, which says the provider is not registered and
names the providers that are.
ProviderRegistry.builtin no longer includes Anthropic. If you used builtin to avoid
classpath discovery (a shaded fat jar, typically), add the module explicitly:
1
given ProviderRegistry = ProviderRegistry.ofModules(new Llm4sAnthropicModule) // `builtin` was removed later in slice 5
Source breaks
Three names could not keep their fully-qualified path, because they were members of objects that stay in core:
ProviderModelListers.Anthropicis nowAnthropicModelLister, in the same package (org.llm4s.config).AnthropicProvider.modelListerreturns it, so code that reached the lister through the descriptor is unaffected.DefaultConfig.DEFAULT_ANTHROPIC_BASE_URLis nowAnthropicConfig.DEFAULT_BASE_URL, the same placeGeminiConfig,DeepSeekConfigandMistralConfigkeep theirs. The value is unchanged.ConfigKeys.ANTHROPIC_API_KEYandConfigKeys.ANTHROPIC_BASE_URLare nowAnthropicConfigKeys.ANTHROPIC_API_KEYandAnthropicConfigKeys.ANTHROPIC_BASE_URL, in the same package, asConfigKeys.OLLAMA_*becameOllamaConfigKeys. The strings are unchanged.
What did not change
Every configuration key and environment variable: llm4s.providers.<name> with
provider = "anthropic", apiKey and the optional baseUrl.
What names Anthropic without depending on its client stays in core and answers the same with or
without llm4s-anthropic: ToolRegistry.getToolDefinitionsSafe("anthropic"), the
anthropic/... entries in the embedded model registry data, the sk-ant- secret pattern,
config-policy allow-lists, and AnthropicStreamingHandler - the SDK-free SSE parser behind
StreamingResponseHandler.forProvider("anthropic"), which AnthropicClient does not use.
Slice 5: llm4s-gemini
The second provider module of slice 5 (#1132),
carrying both of Google’s chat providers: the Gemini API (provider = "gemini", alias
"google") and Vertex AI (provider = "vertexai", alias "vertex"). It is in the build but
not yet in a release; 0.4.1 and earlier still ship both inside llm4s-core.
Why Vertex AI is in the same module
VertexAIClient only calls Google’s publishers/google models - Gemini - using the same JSON
request and response format as GeminiClient. The two differ in endpoint (Vertex is scoped to
a GCP project and region on aiplatform.googleapis.com) and in authentication (Vertex uses
OAuth2, implemented by VertexAIAuthProvider without a Google SDK), not in dependencies. So
bundling them costs a Gemini-API user nothing, while splitting Vertex AI out later would be a
breaking move for its users; bundling now is the direction that stays safe.
What moved
| Code | Now in |
|---|---|
GeminiClient, GeminiProvider (org.llm4s.llmconnect.provider) |
llm4s-gemini |
VertexAIClient, VertexAIProvider, VertexAIAuthProvider (org.llm4s.llmconnect.provider) |
llm4s-gemini |
GeminiConfig, VertexAIConfig (org.llm4s.llmconnect.config) |
llm4s-gemini |
ProviderModelListers.Gemini → GeminiModelLister (org.llm4s.config) |
llm4s-gemini |
DefaultConfig.DEFAULT_GEMINI_BASE_URL → GeminiConfig.DEFAULT_BASE_URL |
llm4s-gemini |
DefaultConfig.DEFAULT_VERTEXAI_LOCATION → VertexAIConfig.DEFAULT_LOCATION (already existed) |
llm4s-gemini |
the commented gemini-main example in reference.conf |
llm4s-gemini’s reference.conf, with a vertexai-main example beside it |
Package names are unchanged, so import org.llm4s.llmconnect.config.GeminiConfig keeps
working once the dependency is added:
1
libraryDependencies += "org.llm4s" %% "llm4s-gemini" % version
Registration is the dependency
llm4s-gemini declares Llm4sGeminiModule in its META-INF/services, so
ProviderRegistry.default finds it and provider = "gemini", "google", "vertexai" and
"vertex" resolve as before. Without the dependency they fail with the registry’s error, which
says the provider is not registered and names the providers that are.
ProviderRegistry.builtin no longer includes Gemini or Vertex AI. If you used builtin to
avoid classpath discovery (a shaded fat jar, typically), add the module explicitly:
1
given ProviderRegistry = ProviderRegistry.ofModules(new Llm4sGeminiModule) // `builtin` was removed later in slice 5
Source breaks
Three names could not keep their fully-qualified path, because they were members of objects that stay in core:
ProviderModelListers.Geminiis nowGeminiModelLister, in the same package (org.llm4s.config).GeminiProvider.modelListerreturns it, so code that reached the lister through the descriptor is unaffected.DefaultConfig.DEFAULT_GEMINI_BASE_URLis nowGeminiConfig.DEFAULT_BASE_URL, the same placeDeepSeekConfig,CohereConfigandMistralConfigkeep theirs. The value is unchanged.DefaultConfig.DEFAULT_VERTEXAI_LOCATIONis removed; useVertexAIConfig.DEFAULT_LOCATION, which already held the same"us-central1".
What did not change
Every configuration key and environment variable: llm4s.providers.<name> with
provider = "gemini" or "vertexai", the Vertex AI reading of endpoint (GCP project id),
organization (region) and apiKey (credential file path), and GOOGLE_APPLICATION_CREDENTIALS
for Vertex AI authentication. (endpoint and organization have since become Vertex AI’s
project and location; see
the note above.)
Strings that name Gemini without depending on its client stay in core and answer the same with
or without llm4s-gemini: ToolRegistry.getToolDefinitionsSafe("gemini"), the gemini/...
entries in the embedded model registry data, and config-policy allow-lists.
Slice 4 (close-out): the last closed provider list, and fromValues stops throwing
The last items deferred from slice 4 (#1131).
Both are source breaks, taken now because the API is not yet frozen; neither has a
deprecated shim, following the precedent of ProviderKind in PR 1. It is in the build but not
yet in a release, so nothing here affects 0.4.1 or earlier.
org.llm4s.rag.EmbeddingProvider is gone
llm4s-rag kept its own closed list of embedding providers - EmbeddingProvider.OpenAI,
Voyage and Ollama - duplicating what the ProviderRegistry has known since PR 4. It was
wrong in both directions: an embedding provider from its own module could not be named through
it, and it named ollama whether or not llm4s-ollama was on the classpath. It also shared its
simple name with org.llm4s.llmconnect.provider.EmbeddingProvider, the embedding client trait.
RAGConfig now names the provider by id, the same id as in EMBEDDING_MODEL=<id>/<model>, and
RAG.build resolves it through the registry:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
// Before
import org.llm4s.rag.{ EmbeddingProvider, RAG }
RAG.builder()
.withEmbeddings(EmbeddingProvider.OpenAI, "text-embedding-3-large")
EmbeddingProvider.fromString(name).toRight(...) // to turn a configured name into one
// After
import org.llm4s.rag.RAG
RAG.builder()
.withEmbeddings("openai", "text-embedding-3-large")
RAG.builder().withEmbeddings(name) // any registered id or alias; no conversion
RAGConfig.embeddingProvider is a ProviderId, so config.embeddingProvider.name becomes
config.embeddingProvider.asString. RAG.build and RAGConfig#build take an implicit
ProviderRegistry, resolved to ProviderRegistry.default when none is in scope, so existing
call sites compile unchanged and an application with its own registry reaches the providers in
it. EmbeddingProvider.values has no replacement in llm4s-rag: the providers are
summon[ProviderRegistry].embeddingIds.
Behaviour changes
- An id that is not registered fails
RAG.buildwith the registry’s own “Embedding provider ‘…’ is not registered” error, naming the ids that are - wherefromStringreturnedNoneand every caller wrote its own message. The resolver is asked for the provider’s canonical id, so an alias such asvoyageaiarrives asvoyage. - The model default moved from
llm4s-raginto the provider.RAGused to pick a model per provider from its own table, and ignore the model in theEmbeddingProviderConfigthe resolver returned.RAGConfig()still defaults toopenai/text-embedding-3-small, butwithEmbeddings(provider)without a model now clears any model set earlier and uses, in order: the resolved config’s model, then the provider’sconfigSpec.defaultModel. WithLlm4sConfig.embeddings()as the resolver, that is the model you configured. A provider with neither fails the build naming the provider, rather than guessing. - Dimensions come from the provider’s
dimensionsOf(model)when not set explicitly, rather than from a second table inllm4s-rag. A model its provider does not declare keeps the old fallback of 1536.
fromValues returns Result
Every ProviderConfig subtype’s fromValues factory - OpenAIConfig, AzureConfig,
AnthropicConfig, ZaiConfig, GeminiConfig, DeepSeekConfig, CohereConfig,
MistralConfig, VertexAIConfig and OllamaConfig - validated its arguments with
require(...), so a blank API key threw IllegalArgumentException out of a library whose rule
is that errors are values. They now return Result[XConfig], and a blank credential or endpoint
is a ConfigurationError carrying the same message as before ("OpenAI apiKey must be
non-empty") with the field in missingKeys.
1
2
3
4
5
6
7
// Before
val config: OpenAIConfig = OpenAIConfig.fromValues("gpt-4o", apiKey, None, baseUrl)
val client = LLMConnect.getClient(config)
// After
val client: Result[LLMClient] =
OpenAIConfig.fromValues("gpt-4o", apiKey, None, baseUrl).flatMap(LLMConnect.getClient(_))
A ProviderDescriptor.buildConfig that returned fromValues from a for’s yield, or
through .map, binds it as a generator or uses .flatMap instead:
1
2
3
4
5
6
7
8
9
10
11
12
// Before
for
apiKey <- ProviderDescriptor.requireApiKey(providerName, section)
baseUrl <- ProviderDescriptor.resolveBaseUrl(providerName, section, configSpec)
yield AcmeConfig.fromValues(section.model.asString, apiKey, baseUrl)
// After
for
apiKey <- ProviderDescriptor.requireApiKey(providerName, section)
baseUrl <- ProviderDescriptor.resolveBaseUrl(providerName, section, configSpec)
config <- AcmeConfig.fromValues(section.model.asString, apiKey, baseUrl)
yield config
Code that caught the exception - Try(OpenAIConfig.fromValues(...)).toEither, or a test’s
an[IllegalArgumentException] should be thrownBy - matches on the Left instead. Nothing
reachable from configuration changes: Llm4sConfig and the provider descriptors already
rejected a missing key before calling fromValues, and now propagate its Left rather than
letting a blank one throw.
Slice 5: llm4s-ollama
The first provider module of slice 5 (#1132).
Ollama goes first because it has the smallest client, no vendor SDK, and a live @Ollama
integration tier - so the carve is checked against a real server rather than mocks. It is in
the build but not yet in a release; 0.4.1 and earlier still ship Ollama inside llm4s-core.
What moved
| Code | Now in |
|---|---|
OllamaClient, OllamaProvider, OllamaEmbeddingProvider (org.llm4s.llmconnect.provider) |
llm4s-ollama |
OllamaConfig (org.llm4s.llmconnect.config) |
llm4s-ollama |
ProviderModelListers.Ollama → OllamaModelLister (org.llm4s.config) |
llm4s-ollama |
ConfigKeys.OLLAMA_* → OllamaConfigKeys.OLLAMA_* (org.llm4s.config) |
llm4s-ollama |
the llm4s.embeddings.ollama reference.conf block |
llm4s-ollama’s reference.conf |
Package names are unchanged, so import org.llm4s.llmconnect.config.OllamaConfig keeps
working once the dependency is added:
1
libraryDependencies += "org.llm4s" %% "llm4s-ollama" % version
Registration is the dependency
llm4s-ollama declares Llm4sOllamaModule in its META-INF/services, so
ProviderRegistry.default finds it and provider = "ollama" and
EMBEDDING_MODEL=ollama/<model> resolve as before. Without the dependency both fail with the
registry’s error, which says ollama is not registered and names the providers that are.
ProviderRegistry.builtin no longer includes Ollama, because core no longer ships it. If you
used builtin to avoid classpath discovery (a shaded fat jar, typically), add the module
explicitly:
1
given ProviderRegistry = ProviderRegistry.ofModules(new Llm4sOllamaModule) // `builtin` was removed later in slice 5
Source breaks
Two names could not keep their fully-qualified path, because they were members of objects that stay in core:
ProviderModelListers.Ollamais nowOllamaModelLister, in the same package (org.llm4s.config).OllamaProvider.modelListerreturns it, so code that reached the lister through the descriptor is unaffected.ConfigKeys.OLLAMA_BASE_URL,OLLAMA_EMBEDDING_BASE_URLandOLLAMA_EMBEDDING_MODELare now onOllamaConfigKeys, also inorg.llm4s.config. The variable names themselves are unchanged.
What did not change
Every configuration key and environment variable: llm4s.providers.<name> with
provider = "ollama", llm4s.embeddings.ollama.*, OLLAMA_EMBEDDING_BASE_URL and
OLLAMA_EMBEDDING_MODEL. The reference.conf block moved rather than changed; HOCON merges
reference files across jars, so the keys exist exactly when the provider does.
TokenizerMapping still recognises the ollama/ model-name prefix. It is a naming
convention on model strings rather than a reference to the provider, and it gives the same
answer whether or not llm4s-ollama is on the classpath.
Slice 4 follow-up: embedding dimensions move into the provider
One of the two items deferred from slice 4 (#1131),
and the precursor to carving llm4s-ollama (#1132).
ModelDimensionRegistry was the last central provider list on the embedding side: a map in
llm4s-core covering openai, voyage and local. It had no ollama entry, so the documented
1
EMBEDDING_MODEL=ollama/nomic-embed-text
failed at Llm4sConfig.textEmbeddingModel() with Unknown model 'nomic-embed-text' for
provider 'ollama', and RAGASFactory.fromConfigs papered over the same gap with
.getOrElse(1536) for a model that is 768-dimensional.
Declaring dimensions
An embedding provider now declares the dimensions of the models it knows, alongside its
configSpec:
1
2
3
4
5
6
object JinaEmbeddings extends EmbeddingProviderDescriptor:
val id = ProviderId("jina")
override val modelDimensions = Map(
"jina-embeddings-v3" -> 1024
)
A provider whose model names have variants overrides dimensionsOf(model) instead - Ollama
folds a :latest tag onto the untagged name, and no other tag, because other tags of one model
can differ in size. A model a provider does not declare still embeds; only a caller that needs
its dimensionality up front is told it is unknown.
Source-compatible signature changes
ModelDimensionRegistry.getDimension, RAGASFactory.fromConfigs and
RAGASFactory.basicFromConfigs take an implicit ProviderRegistry, resolved to
ProviderRegistry.default when none is in scope. Existing call sites compile unchanged; a
caller with its own registry now reaches the providers in it.
ModelDimensionRegistry.localDimension(model) answers the local non-text encoders
(openclip-vit-b32, wav2vec2-base, timesformer-base), which have no descriptor because
nothing can be configured with them. getDimension("local", ...) still works.
Behaviour changes
RAGASFactory.fromConfigsandbasicFromConfigsreturn the lookup’sLeftfor an embedding model its provider does not declare, instead of assuming 1536 dimensions. Build theEmbeddingModelConfigyourself and callRAGASFactory.createorbasicfor such a model.getDimensionfor an unregistered provider returns the registry’s “not registered” error, naming the embedding providers that are, rather than “Unknown model”.voyage-3-largeresolves to 1024, its default output size, not 1536.
Slice 4 (PR 5): embedding config moves into the provider
The fifth slice 4 change (#1131), and the follow-up PR 4 named. PR 4 made an embedding provider resolvable from its own module; its configuration was still core’s business:
1
2
3
4
5
final private case class EmbeddingsOllamaSection(apiKey: …, baseUrl: …, model: …)
implicit private val embeddingsOllamaSectionReader = …
private val DefaultOllamaEmbeddingBaseUrl = "http://localhost:11434"
private def buildOllamaEmbeddings(…) = …
// plus two `match` arms
— a typed case class, a PureConfig reader, a default, a builder and two dispatch arms per provider. A third-party embedding provider could be registered and then had nothing to be configured with.
Nothing changes for users. EMBEDDING_MODEL, EMBEDDING_PROVIDER, OPENAI_API_KEY,
VOYAGE_API_KEY, OLLAMA_EMBEDDING_BASE_URL and every llm4s.embeddings.<id> key behave
exactly as before.
The section shape is core’s; what it means is the provider’s
llm4s-core now parses one uniform shape - apiKey, baseUrl, model - for whichever
provider was selected, and hands it to the descriptor:
1
def buildConfig(section: EmbeddingProviderSection, modelOverride: Option[String]): Result[EmbeddingProviderConfig]
Everything it needs arrives in section, already typed. A descriptor reads no configuration
itself: raw config access stays in org.llm4s.config, which is the boundary AGENTS.md sets.
Most providers never implement it. Declaring an EmbeddingConfigSpec is enough, and the
default implementation resolves the three fields against it:
1
2
3
4
5
6
7
8
object JinaEmbeddings extends EmbeddingProviderDescriptor:
val id = ProviderId("jina")
override val configSpec = EmbeddingConfigSpec(
requiresApiKey = true,
defaultBaseUrl = Some("https://api.jina.ai/v1"),
apiKeyEnv = Some("JINA_API_KEY") // named in the error when it is missing
)
Defaults are code, environment bindings are HOCON
They used to be both. reference.conf said baseUrl = "http://localhost:11434" and
EmbeddingsConfigLoader said DefaultOllamaEmbeddingBaseUrl, with nothing keeping them in
step. The default now lives only in the descriptor’s EmbeddingConfigSpec, and each
provider’s reference.conf block is reduced to the environment variables it binds:
1
2
3
4
ollama {
baseUrl = ${?OLLAMA_EMBEDDING_BASE_URL}
model = ${?OLLAMA_EMBEDDING_MODEL}
}
That block is keyed by provider id, so it travels with the provider when the provider moves
to its own module - HOCON merges these across jars. ProviderId canonicalises (trim, lowercase)
but does not restrict, so an id containing a dot is legal and must be quoted as a single HOCON
key; EmbeddingConfigSpec.sectionPath / fieldPath build the path that way, and the paths
named in errors are the paths to write:
1
llm4s.embeddings."acme.embeddings" { apiKey = ${?ACME_API_KEY} }
(Chat config cannot do this: it is keyed by the user’s instance name, which is why
ProviderConfigSpec.defaultBaseUrl is code and says so.)
A key that lives somewhere else
Superseded:
apiKeyPathwas removed in favour of the sharedllm4s.credentials.<id>.apiKey- see Vendor credentials.apiKeyEnvis now aSeq[String].
A provider whose key is kept outside its own llm4s.embeddings.<id> section declares where to
look instead of the loader special-casing it, and that declaration is also what makes the error
name the place the key is really set:
1
2
3
4
5
6
override val configSpec = EmbeddingConfigSpec(
requiresApiKey = true,
apiKeyPath = Some("llm4s.acme.apiKey"),
apiKeyEnv = Some("ACME_API_KEY"),
…
)
Missing acme embeddings apiKey (llm4s.acme.apiKey / ACME_API_KEY)
OpenAI was the case this was written for, reading llm4s.openai.apiKey; it no longer uses it
(its key is its own llm4s.embeddings.openai.apiKey), and no provider in this repository does.
apiKeyPath is a declaration, not a read: EmbeddingsConfigLoader resolves it and hands
the value back in the section before calling buildConfig. The provider owns the knowledge of
where its key lives; org.llm4s.config keeps sole ownership of reading it.
The embedding config entry points take the registry
1
2
3
def embeddings()(using ProviderRegistry): Result[(String, EmbeddingProviderConfig)]
def loadTextEmbeddingModel()(using ProviderRegistry): Result[TextEmbeddingModelSettings]
def textEmbeddingModel()(using ProviderRegistry): Result[TextEmbeddingModelSettings]
Binary-incompatible, source-compatible, as in PRs 3 and 4. Without it an application’s own
registry could not reach the loader, and a provider it registered explicitly would resolve for
EmbeddingClient.from but not for its configuration.
All three resolve through the same registry, so a provider configurable by one is configurable
by all of them: textEmbeddingModel now goes through embeddings rather than calling the
loader a second time.
Error messages
Unknown providers now produce the registry’s message, which names what is registered and the
scan that found it. Missing-field errors name the config path and the environment variable the
descriptor declared. The provider is named by its canonical id (openai, not OpenAI) - the
spelling that appears in config.
Slice 4 (PR 4): embedding providers join the SPI
The fourth slice 4 change (#1131), and the last
unchecked item on that issue’s list. PRs 2 and 3 made a chat provider self-describing and
discoverable. EmbeddingClient.from was still the shape the slice exists to delete:
1
2
3
4
5
provider.toLowerCase match
case "openai" => Right(new EmbeddingClient(OpenAIEmbeddingProvider.fromConfig(cfg)))
case "voyage" => ...
case "ollama" => ...
case other => Left(EmbeddingError(...))
so an embedding provider was an edit to llm4s-core no matter where its code lived. It is now
resolved through the same ProviderRegistry.
Declaring an embedding provider
EmbeddingProviderDescriptor is the embedding counterpart of ProviderDescriptor, and
Llm4sProviderModule gained a second list:
1
2
3
4
5
6
7
8
object JinaEmbeddings extends EmbeddingProviderDescriptor:
val id = ProviderId("jina")
def build(config: EmbeddingProviderConfig): Result[EmbeddingProvider] =
Right(JinaEmbeddingProvider.fromConfig(config))
final class JinaProviderModule extends Llm4sProviderModule:
override def embeddingProviders: Seq[EmbeddingProviderDescriptor] = Seq(JinaEmbeddings)
Registration is otherwise identical to PR 3 - the same META-INF/services file, the same
ServiceLoader scan, the same escape hatches. A module declares whichever halves it has; both
default to empty.
If your module delegates, forward both lists. Every
Llm4sProviderModulemember defaults toNil, so a module that forwards onlychatProviderscontributes no embedding providers and fails silently rather than at compile time.
Why a separate trait, not a method on ProviderDescriptor
The two provider sets overlap without either containing the other: OpenAI and Ollama supply a
chat client and an embedding provider, Voyage supplies only embeddings, Anthropic only chat.
Folding embeddings into ProviderDescriptor would force an embedding-only provider to implement
buildConfig and buildClient only to fail them.
The ids therefore live in two namespaces, and the same id can appear in both — ollama names
a chat client and an embedding provider that share nothing but a base URL. ids and
embeddingIds list them separately, and a provider that supplies no embeddings fails as such:
Embedding provider ‘anthropic’ (from llm4s.embeddings.model) is not registered. Registered embedding providers: ollama, openai, voyage. If you expected ‘anthropic’, add the dependency that supplies it, or register it explicitly with ProviderRegistry.ofEmbeddings(…).
Each half names the registration call that accepts its own descriptor type - of for chat,
ofEmbeddings for embeddings - because following the other one is a compile error.
EmbeddingClient.from takes the registry
1
2
3
4
def from(provider: String, cfg: EmbeddingProviderConfig)(using
ModelRegistryService,
ProviderRegistry
): Result[EmbeddingClient]
Binary-incompatible, source-compatible: existing call sites resolve ProviderRegistry.default
through the companion’s given, exactly as the Llm4sConfig methods did in PR 3. An application
that registers its own passes it:
1
2
given ProviderRegistry = ProviderRegistry.default.withEmbeddingProvider(JinaEmbeddings)
EmbeddingClient.from("jina", cfg)
ProviderRegistry gained findEmbedding, resolveEmbedding, embeddingIds,
canonicalEmbeddingId, withEmbeddingProvider and ofEmbeddings; ProviderModuleReport gained
embeddingProviderIds. The unknown-provider failure is still an EmbeddingError with code
400, now carrying the registry’s diagnostics as its message.
What did not change
EmbeddingProvider itself, EmbeddingProviderConfig, and each provider’s fromConfig are
untouched - OpenAIEmbeddingProvider.fromConfig(cfg) still works and is still the direct route.
The three built-in objects simply are their own descriptors now.
Embedding configuration is not part of this change: llm4s.embeddings still has typed
openai / voyage / ollama sections in EmbeddingsConfigLoader, so a third-party embedding
provider is reachable through EmbeddingClient.from but still needs its config built by the
application. Moving config binding into the descriptor, as PR 2 did for chat, is the follow-up.
Slice 4 (PR 3): providers are discovered on the classpath
The third slice 4 change (#1131). PR 2 made a
provider a ProviderDescriptor that registers itself; this removes the last manual step. A
provider module on the classpath is now found without any registration code at the call site -
adding a provider is adding a dependency.
Declaring a provider module
Ship a META-INF/services/org.llm4s.llmconnect.spi.Llm4sProviderModule naming an implementation:
1
com.example.llm4s.BedrockProviderModule
1
2
3
4
5
6
// Must be a `class` with a public no-arg constructor, not an `object`:
// ServiceLoader instantiates the named class, and a Scala `object` exposes its
// instance as a MODULE$ field instead. (This is also what GraalVM native-image
// needs, via its ServiceLoaderFeature.)
final class BedrockProviderModule extends Llm4sProviderModule:
override def chatProviders: Seq[ProviderDescriptor] = Seq(BedrockProvider)
That is the whole registration. ProviderRegistry.default - what every Llm4sConfig and
LLMConnect call uses when the caller supplies no registry - is now
ProviderRegistry.discover(), computed once on first use.
llm4s-core declared its own providers the same way, through
org.llm4s.llmconnect.provider.BuiltinProviderModule, with no special case: they were
discovered exactly as a third-party module is. (Since slice 5 core ships no provider, and that
module is gone.)
One broken jar cannot take out the others
java.util.ServiceLoader’s iterator throws ServiceConfigurationError for an entry it cannot
load, and the for-comprehension you would naturally write over it propagates the first such
error and abandons every remaining provider. discover drives the iterator by hand and guards
each step, so an unusable entry becomes a recorded failure and the scan continues:
1
2
3
4
5
6
val registry = ProviderRegistry.discover()
registry.report.failures.foreach(f => println(f.detail))
println(registry.report.describe)
// Discovery scanned 2 modules; 1 failed: loading a provider module failed: ...
// - org.llm4s.llmconnect.provider.BuiltinProviderModule [file:/.../llm4s-core.jar]: openai, openrouter, ...
// ! loading a provider module failed: ... Provider com.example.Missing not found
Failures are logged at WARN as they happen, and the scan summary is appended to the “provider is not registered” error, because the two failure modes that are otherwise invisible are a dependency that was never added and a fat jar whose services files were dropped:
Provider ‘bedrock’ (from llm4s.providers.my-bedrock.provider) is not registered. Registered providers: anthropic, azure, … If you expected ‘bedrock’, add the dependency that supplies it, or register it explicitly with ProviderRegistry.of(…). Discovery scanned 1 module; 0 failed.
Fat jars
Shading tools default to overwriting same-named resources, which silently discards every services file but one. Configure them to concatenate:
1
2
3
4
5
// sbt-assembly
assembly / assemblyMergeStrategy := {
case PathList("META-INF", "services", _*) => MergeStrategy.filterDistinctLines
case other => (assembly / assemblyMergeStrategy).value(other)
}
1
2
<!-- maven-shade -->
<transformer implementation="org.apache.maven.plugins.shade.resource.ServicesResourceTransformer"/>
If you cannot, register explicitly - this is what the escape hatch is for:
1
2
val registry = ProviderRegistry.ofModules(new Llm4sOpenAIModule).withProvider(BedrockProvider)
LLMConnect.getClient(config)(using registry)
ProviderRegistry.ofModules builds a registry from the modules you name, with no classpath scan
at all. (This section first suggested ProviderRegistry.builtin, the providers compiled into
llm4s-core; it was removed when the last of them left core.)
Llm4sConfig takes the registry
Every Llm4sConfig method that reads llm4s.providers now takes an implicit
ProviderRegistry: provider, providerConfigs (both), providers, defaultProviderName,
defaultProvider, listModels (both), and providerFrom. Existing call sites are unchanged -
the companion supplies ProviderRegistry.default - and a caller who wants a different set of
providers passes one:
1
2
given ProviderRegistry = ProviderRegistry.default.withProvider(MyProvider)
val config = Llm4sConfig.provider("my-provider") // now resolvable
This is a binary-incompatible change to those signatures, and source-compatible.
What did not change
ProviderDescriptor, ProviderConfigSpec, ProviderFeatures and Llm4sProviderModule are as
PR 2 shipped them. ProviderRegistry.of, ofModules, withProvider and withModule behave as
before; registries built that way report discovered = false and carry no scan summary.
Slice 4 (PR 2): the provider registration SPI
The second slice 4 change (#1131). PR 1 removed the
two structures that made an out-of-module provider impossible - a closed enum and a sealed
trait. This one builds the extension point on top: a provider is now a value that describes
itself, and everything that used to enumerate providers looks them up instead.
Every provider still ships inside llm4s-core; what changed is that none of them is wired in
by hand any more. Splitting them into their own artifacts is slice 5
(#1132).
Adding a provider
Before, adding one chat provider meant editing roughly eight shared files - the closed enum,
the sealed config file, two match expressions in LLMConnect, the loader’s dispatch, a
validator object, a capabilities object and the capabilities registry. That surface is why 13
open provider PRs all conflict with each other.
Now it is one file:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
object BedrockProvider extends ProviderDescriptor:
val id: ProviderId = ProviderId("bedrock")
val configSpec: ProviderConfigSpec =
ProviderConfigSpec(requiresApiKey = true, requiresEndpoint = true)
def buildConfig(providerName: String, section: NamedProviderConfig)(using
ContextWindowResolver
): Result[ProviderConfig] =
for
apiKey <- ProviderDescriptor.requireApiKey(providerName, section)
endpoint <- ProviderDescriptor.requireField(
providerName, "endpoint", section.endpoint, "llm4s.providers.<name>.endpoint")
yield BedrockConfig.fromValues(section.model.asString, apiKey, endpoint)
def buildClient(config: ProviderConfig, options: LlmClientOptions)(using
ModelRegistryService
): Result[LLMClient] =
ProviderDescriptor
.expectConfig[BedrockConfig](id, config)
.flatMap(BedrockClient(_, options.metrics, options.exchangeLogging))
plus a ProviderConfig implementation, a client, and one registration:
1
2
val registry = ProviderRegistry.default.withProvider(BedrockProvider)
LLMConnect.getClient(config)(using registry)
Classpath discovery - registering by adding a dependency, with no code at all - is PR 3.
Llm4sProviderModule is the service type it will discover; it is already here so that a module
can group its providers today.
What is new
All in org.llm4s.llmconnect.spi:
| Type | Purpose |
|---|---|
ProviderDescriptor |
one provider: id, aliases, config shape, features, model lister, and the two builders |
ProviderConfigSpec |
which section fields the provider requires, its default base URL, and the text shown when a field is missing |
ProviderFeatures |
what the client actually implements, declared statically (Cohere and Mistral declared streaming = false until slice 5 - #925) |
ProviderRegistry |
an immutable set of descriptors; of, withProvider, withModule, and lookup that returns Result |
Llm4sProviderModule |
the unit of registration - one module, several providers |
ProviderRegistry is resolved through a using clause with a default given in its companion, so
existing call sites are unchanged and a caller who wants a different set passes one:
1
2
LLMConnect.getClient(config) // ProviderRegistry.default
LLMConnect.getClient(config)(using myRegistry) // only the providers you registered
What was deleted
All of these were private[llm4s], so this costs users nothing:
| Deleted | Replaced by |
|---|---|
config.ProviderCapabilities (trait + 12 objects) |
ProviderDescriptor |
config.ProviderCapabilitiesRegistry |
ProviderRegistry |
config.NamedProviderValidator (trait) and NamedProviderValidators (12 objects) |
ProviderConfigSpec + one generic NamedProviderSectionValidator |
the twelve-branch match in NamedProviderLoader |
descriptor.buildConfig |
the two match expressions in LLMConnect |
descriptor.buildClient |
the hard-coded "google"/"vertex" alias fold in NamedProviderConfigNormalizer |
ProviderDescriptor.aliases |
Error messages for missing fields are unchanged; they are now generated from the spec rather than written out per provider.
Source breaks
-
ReliableProviders: seven per-provider factories collapse towrap.1 2 3 4 5 6 7
// Before - covered 7 of 12 providers, and no provider from another module ReliableProviders.openai(config, ReliabilityConfig.aggressive) ReliableProviders.anthropic(config) // After - covers every registered provider ReliableProviders.wrap(config, ReliabilityConfig.aggressive) ReliableProviders.wrap(config)
The missing five were DeepSeek, Cohere, Mistral, Requesty and Vertex AI.
wrap(client, providerName, ...), for a client you already have, is unchanged. -
OpenAIConfig.providerIdis derived frombaseUrl. It answersopenrouterfor a URL containingopenrouter.aiandopenaiotherwise - which is exactly the routingLLMConnectalready did with a hard-coded check, now stated by the config itself. If you build anOpenAIConfigfor OpenRouter and inspectproviderId, the answer changed fromopenaitoopenrouter; client routing is unchanged. -
ProviderModelListeris now public, along withProviderModelListersand its newopenAICompatible(provider, defaultBaseUrl, modelsPath)factory - a provider module needs to supply a model lister, and most providers serve the OpenAI/modelsshape. The five near-identical per-provider lister objects became calls to that factory. -
ProviderResultOpsandProviderExchangeRecorderare now public (they wereprivate[provider]). A provider client outsidellm4s-coreneeds both.
What did not change
Llm4sConfig’s signatures, NamedProviderLoader’s results, DiscoveredModel, every
ProviderConfig subtype’s fields, and the llm4s.providers.* config format. A configuration
that worked before works now, including provider = "google" and provider = "vertex".
ProviderConfig.fromValues still uses require(...), which throws rather than returning a
Left. Converting it is a behaviour change (throw → Left) that deserves its own note, and
embeddings (EmbeddingClient.from, EmbeddingsConfigLoader’s fixed-arity reader,
ModelDimensionRegistry, and the duplicate org.llm4s.rag.EmbeddingProvider name ADT) are still
on the old dispatch. Both are tracked under #1131.
Since resolved. The embedding entry points moved onto the registry in PR 4, PR 5 and the dimensions follow-up above; the
org.llm4s.rag.EmbeddingProviderADT was removed, andfromValuesconverted toResult, in the slice 4 close-out.
Slice 4 (PR 1): ProviderKind becomes ProviderId, ProviderConfig opens up
The first of the slice 4 changes (#1131). No SPI
yet - this only removes the two things that make a provider impossible to supply from outside
llm4s-core: a closed enum and a sealed trait. It is in the build but not yet in a release,
so nothing here affects 0.4.1 or earlier.
This is a clean break: there is no deprecated ProviderKind shim. A shim would have kept a
closed list of twelve providers inside the very module whose purpose is to remove it, and it
would only have half-worked - case ProviderKind.OpenAI => and def f(k: ProviderKind) would
still compile, while .values, .ordinal, .fromOrdinal and exhaustivity would not. A clear
compile error beats a partly-working deprecated type.
ProviderKind → ProviderId
1
2
3
4
5
6
7
8
9
// Before
enum ProviderKind:
case OpenAI; case Anthropic; /* ... ten more */
// After
opaque type ProviderId = String
object ProviderId:
def apply(raw: String): ProviderId = raw.trim.toLowerCase(Locale.ROOT) // canonicalises
extension (id: ProviderId) def asString: String = id
ProviderId is an open vocabulary, not an enumeration. Any string names a provider; whether
that provider can be resolved is answered at resolution time, by whatever is on the classpath -
which is the whole point. It stays opaque over String, so Option[ProviderId] and
Map[ProviderName, ProviderId] still do not box.
| Before | After |
|---|---|
ProviderKind.OpenAI |
ProviderId("openai") |
ProviderKind.fromString(s) / fromName(s), returning Option |
ProviderId(s), total |
kind.name |
id.asString |
kind.toString → "OpenAI" |
id.asString → "openai" |
ProviderKind.all |
no replacement - ask the thing that resolves providers, not the type |
ProviderKind.values / .ordinal / .fromOrdinal / .productPrefix |
no replacement |
exhaustive match on ProviderKind |
match on id.asString, with a default branch |
asString lives in object ProviderId rather than beside ModelName.asString and friends,
because every newtype in ProviderModelTypes erases to String and a second asString at that
level would be a double definition after erasure. Companion-scoped extensions resolve through the
opaque type’s implicit scope, so id.asString still needs no extra import.
toString changed value. ProviderKind.OpenAI.toString was "OpenAI"; a ProviderId is
its canonical lowercase spelling, so it prints "openai". If you interpolated a provider into
log or error text, expect the case to change. Uppercase derivations still work:
providerId.asString.toUpperCase is "OPENAI", as providerKind.toString.toUpperCase was.
ProviderConfig is no longer sealed, and describes itself
In Scala 3 sealed restricts extension to the same file, so all ten provider configs were
stuck in one 756-line file - not merely in the same jar. ProviderConfig is now a plain trait
with three new members:
1
2
3
4
5
trait ProviderConfig:
def providerId: ProviderId // replaces `val provider: ProviderKind`
def endpointUrl: Option[String] // the endpoint this config will contact
def withModel(model: String): ProviderConfig
// model, contextWindow, reserveCompletion unchanged
1
2
3
4
// Before
config.provider == ProviderKind.OpenAI
// After
config.providerId == ProviderId("openai")
The three additions exist so that code describing a config does not have to know the set of
subtypes. Four exhaustive matches were deleted rather than moved by using them:
ConfigPolicyEngine.providerName, ConfigPolicyEngine.baseUrlOrEndpoint,
PrometheusMetricsExample’s provider-name match, and ProviderSetupRuntime.overrideModel. If
you have a match on ProviderConfig, that is the migration: reach for providerId,
endpointUrl or withModel first, and only keep the match if you genuinely need
provider-specific fields.
Losing sealed also means an exhaustive match on ProviderConfig now compiles with a
warning - and fails for anyone building with -Werror, as this repo does. Add a default
branch, or annotate the scrutinee (config: @unchecked) if you have deliberately accepted the
risk.
One behaviour change falls out of this: ConfigPolicyEngine.baseUrlOrEndpoint used to return
None for VertexAIConfig, because the old match had no case for it. It now returns
Some(computedBaseUrl). A requiredBaseUrlPattern policy that silently reported “no
endpoint/baseUrl found” for Vertex AI will now actually check the URL.
Unknown provider ids are no longer rejected while parsing
NamedProviderConfigNormalizer used to fail on an unrecognised provider string with
"Configured provider 'x' has unknown provider 'moonbeam'". It now produces a ProviderId
unconditionally; only resolution fails, with an error naming what is registered:
1
2
3
No provider capabilities registered for provider 'moonbeam'.
Registered providers: anthropic, azure, cohere, deepseek, gemini, mistral, ollama,
openai, openrouter, requesty, vertexai, zai
This is what lets a provider live in a module llm4s-core has never heard of. The accepted
aliases are unchanged - provider = "google" still resolves to gemini, and
provider = "vertex" to vertexai - though that table moves into each provider’s descriptor
when the SPI lands.
Validation error text is now provider-agnostic
1
2
3
4
// Before
Azure OpenAI provider 'my-azure' is missing required fields:
// After
Provider 'my-azure' (provider = azure) is missing required fields:
The per-field guidance underneath is unchanged, including the ${?AZURE_API_KEY} substitution
hint. (Since superseded: the apiKey line now names the variable the provider module binds -
AZURE_OPENAI_API_KEY for Azure - see
Vendor credentials.) Only the leading sentence differs, because it used to be generated from a hard-coded
display name per provider.
Bug fix: provider = "vertexai" now works at all
ProviderKind.VertexAI existed, NamedProviderLoader built a VertexAIConfig from it, and
LLMConnect built a VertexAIClient from that - but Vertex AI was missing from
ProviderCapabilitiesRegistry and had no validator object. Since validation routes through that
registry, every provider = "vertexai" config failed validation outright, so none of the
supporting code was reachable from configuration. Both are now present, and the config path is
covered by tests.
What did not change
ProviderConfig.fromValues’s require(...) calls still throw rather than returning Result;
converting them is a throw-to-Left behaviour change and is deferred to PR 2.
ReliableProviders’ seven per-provider factories are also unchanged here - they collapse to a
single registry-routed wrap in PR 2. NamedProviderLoader, NamedProviderValidator,
ProviderCapabilities, ProviderCapabilitiesRegistry and ProviderModelLister are all
private[llm4s] or private[config], so their reshaping costs users nothing.
Slice 3: llm4s-speech
The last artifact of slice 3 (#1130) and the last
package out of llm4s-core in this slice. It is in the build but not yet in a release, so
nothing here affects 0.4.1 or earlier.
What moved
| Packages | New module |
|---|---|
org.llm4s.speech (and speech/io, processing, stt, tts, util) |
llm4s-speech |
Package names did not change, and there are no source breaks. 16 main and 21 test files moved whole; the only code outside the package that referenced it was a sample.
1
2
3
4
5
6
7
8
// Before
libraryDependencies += "org.llm4s" %% "llm4s-core" % version
// After - only if you use speech-to-text or text-to-speech
libraryDependencies ++= Seq(
"org.llm4s" %% "llm4s-core" % version,
"org.llm4s" %% "llm4s-speech" % version
)
What llm4s-core sheds
Vosk (25 MB) and JNA. com.alphacephei:vosk is imported by exactly one file,
speech/stt/VoskSpeechToText.scala, and until now sat on the classpath of every llm4s-core
user, whether or not they had any use for offline speech recognition. It is the single largest
dependency the carve programme has moved.
Deps.jna moves with it, but it is worth being precise about why, because it is not a second
dependency: Vosk’s own POM already depends on net.java.dev.jna:jna:5.7.0. The explicit
declaration exists to win that version conflict and pull 5.19.1 instead. Dropping it would not
remove JNA - it would silently downgrade it to a release that predates Apple Silicon support.
llm4s-workspace-client also declared both, and used neither; those declarations are removed
here too, so Vosk genuinely leaves the build for everyone who is not doing speech.
The Vosk resolver is deleted, not moved
build.sbt carried the project’s only third-party resolver:
1
resolvers += "Vosk Repository" at "https://alphacephei.com/maven/"
It resolved nothing. Vosk 0.3.45 publishes to Maven Central, which is where every build has
actually been getting it - the local Coursier cache holds vosk-0.3.45.jar under
repo1.maven.org and not a single artifact under alphacephei.com. What it did do was add a
third-party host to the lookup path for every artifact in the build, including llm4s’s own
inter-module jars, which produced a steady trickle of failed requests to alphacephei.com on
every resolve.
So it is removed rather than carried into llm4s-speech, and the build now has no
third-party resolvers at all.
Slice 3 is complete
With this, llm4s-core no longer contains rag, knowledgegraph, agent/memory, mcp,
imagegeneration, imageprocessing or speech. What remains is the agent runtime,
llmconnect, toolapi, config, trace and the provider clients, which
slices 4 to 6 address.
Slice 3: llm4s-image
Part of slice 3 of the module carves tracked in
#1126; the slice is
#1130. It is in the build but not yet in a
release, so nothing here affects 0.4.1 or earlier.
What moved
| Packages | New module |
|---|---|
org.llm4s.imagegeneration |
llm4s-image |
org.llm4s.imageprocessing |
llm4s-image |
Package names did not change, and this carve adds no source breaks of its own. The image
API’s one source break - image formats becoming org.llm4s.media.MediaType - landed earlier,
in llm4s-media, precisely so that this step is a pure file move.
The two packages move together because they are two halves of one subsystem: generate an image, then analyse or convert it. Both are built on the same media vocabulary, and splitting them would leave two artifacts nobody uses apart.
1
2
3
4
5
6
7
8
// Before
libraryDependencies += "org.llm4s" %% "llm4s-core" % version
// After - only if you generate images, or analyse them with a vision model
libraryDependencies ++= Seq(
"org.llm4s" %% "llm4s-core" % version,
"org.llm4s" %% "llm4s-image" % version
)
llm4s-image brings llm4s-media with it, so MediaType is on your classpath either way.
What llm4s-core sheds
No third-party dependency: the image clients are built on Llm4sHttpClient, ujson/upickle
and javax.imageio from the JDK, all of which core keeps for other reasons. What core sheds is
19 source files and about 3,600 lines of a subsystem most users never touch, along with its
edge to llm4s-media - that edge existed only because the image packages were still inside
core, and it leaves with them.
Core’s measured statement coverage rises from 74.05% to 74.89% as a result, since the image code was below core’s average.
One test moved with the code
org.llm4s.async.AsyncErrorHandlingSpec lived in core’s test tree under a package name that
suggests it is about asynchrony in general. Every one of its assertions exercises an image
client - it checks that ImageProcessingClient.analyzeImageAsync and the image generation
clients’ Future { blocking { ... } }.recover { ... } pattern surface thrown exceptions as
Left rather than as a failed Future. It moves to llm4s-image with the code it tests,
keeping its package name.
Had it been left behind it would simply have stopped compiling - but the more useful point is that leaving it would have removed the only coverage of that behaviour from the module that owns it.
Slice 3: llm4s-media
A new module rather than a carve, landed as part of slice 3
(#1130) and ahead of llm4s-image and
llm4s-speech, so those two carves can be pure file moves. It is in the build but not yet in
a release, so nothing here affects 0.4.1 or earlier.
Why
A media type - a MIME string, a canonical file extension, and whether the thing is an image,
audio, video or text - is the one piece of vocabulary every multimodal subsystem needs to
name. Because there was nowhere shared to put it, each grew its own. llm4s-core shipped
three overlapping enumerations of the same handful of image formats:
| Type | Cases | Members |
|---|---|---|
org.llm4s.imagegeneration.ImageFormat |
PNG, JPEG, WEBP | extension, mimeType |
org.llm4s.imageprocessing.ImageFormat |
PNG, JPEG, WEBP, GIF | extension, mimeType |
org.llm4s.imageprocessing.MediaType |
Jpeg, Png, Gif, WebP, Bmp, Tiff | value |
The first two are structurally identical and differ only in package, so a format produced by
image generation could not be handed to image processing without a hand-written conversion.
The third models the same six formats a third way, in the same package as the second. Meanwhile
MediaExtractor in llm4s-rag discriminated on raw MIME prefixes (mimeType.startsWith
("image/")), with no type to name the answer at all.
Carving image and speech out of core without fixing this would have frozen three copies
into three artifacts, where consolidating them later costs a cross-module source break rather
than an in-module one.
What llm4s-media is
Vocabulary only - no I/O, no content sniffing, no third-party dependencies at all:
org.llm4s.media.MediaType-mimeType,extension,category, plusfromExtension,fromPathandfromMimeTypelookups. Sealed; refines intoImageMediaTypeandAudioMediaTypeso an image API can require an image without re-enumerating the cases.org.llm4s.media.MediaCategory-Image,Audio,Video,Text,Application, withfromMimeType.
Deciding what a file actually is from its bytes needs Tika and stays in llm4s-rag; that
code produces a MIME string and resolves it here. That separation is what lets every consumer
depend on llm4s-media without inheriting anything.
Source breaks
This is a source break, taken deliberately ahead of the 1.0 API freeze. All three types
above are replaced by org.llm4s.media.MediaType.
| Before | After |
|---|---|
org.llm4s.imagegeneration.ImageFormat |
org.llm4s.media.ImageMediaType |
org.llm4s.imageprocessing.ImageFormat |
org.llm4s.media.ImageMediaType |
org.llm4s.imageprocessing.MediaType |
org.llm4s.media.MediaType |
ImageFormat.PNG |
MediaType.Png |
ImageFormat.JPEG |
MediaType.Jpeg |
ImageFormat.WEBP |
MediaType.WebP |
ImageFormat.GIF |
MediaType.Gif |
MediaType.Jpeg.value |
MediaType.Jpeg.mimeType |
1
2
3
4
5
6
7
// Before
import org.llm4s.imagegeneration.ImageFormat
val opts = ImageGenerationOptions(format = ImageFormat.PNG)
// After
import org.llm4s.media.MediaType
val opts = ImageGenerationOptions(format = MediaType.Png)
Two lookups changed shape as well. org.llm4s.imageprocessing.MediaType.fromExtension and
.fromPath were total, silently returning JPEG for anything they did not recognise - so a
.txt file reported as an image and the caller could not tell. The replacements return
Option, and callers that genuinely want the old fallback ask for it:
1
2
3
4
5
// Before
val mt = MediaType.fromPath(path) // JPEG if unrecognised
// After
val mt = MediaType.imageFromPath(path).getOrElse(MediaType.Jpeg) // fallback is now visible
AnthropicVisionClient.detectMediaType keeps the old behaviour and its old signature shape -
it still answers JPEG for an unrecognised extension, because that is what the Anthropic API
assumes for an unlabelled image - but now returns an ImageMediaType.
Adding the dependency
Nothing to add today: llm4s-media arrives as a transitive dependency of llm4s-core (via
the image packages, which are still in core) and of llm4s-rag. Declare it directly only if
you name MediaType or MediaCategory in your own signatures.
1
libraryDependencies += "org.llm4s" %% "llm4s-media" % version
Slice 3: llm4s-mcp
Third of the module carves tracked in
#1126; slice 3 is
#1130, which carves three independent
subsystems - mcp, image and speech - one artifact at a time. This note covers mcp;
the other two follow. It is in the build but not yet in a release, so nothing here affects
0.4.1 or earlier.
What moved
| Packages | New module |
|---|---|
org.llm4s.mcp |
llm4s-mcp |
Package names did not change, and there are no source breaks. Nothing outside
org.llm4s.mcp referenced it, so the whole package moved with no facade left behind.
1
2
3
4
5
6
7
8
// Before
libraryDependencies += "org.llm4s" %% "llm4s-core" % version
// After - only if you use the Model Context Protocol client, server or tool registry
libraryDependencies ++= Seq(
"org.llm4s" %% "llm4s-core" % version,
"org.llm4s" %% "llm4s-mcp" % version
)
What llm4s-core sheds
Java-WebSocket. Worth being precise about why, because it is not what
#1130 predicted: MCP does not use WebSockets at
all. Its transports are stdio, HTTP and SSE, built on Llm4sHttpClient and
com.sun.net.httpserver. Deps.websocket was declared on llm4s-core and imported by
nothing in it - the only WebSocket code in the repo is ContainerisedWorkspace in
llm4s-workspace-client, which declares the dependency itself. So core sheds it by dropping a
declaration that was never used, and the dependency does not follow mcp anywhere.
If you depend on llm4s-core and were picking up org.java-websocket transitively, declare
it yourself.
Configuration keys
None. org.llm4s.mcp reads no reference.conf keys and no Llm4sConfig method names a type
that moved.
Slice 2: llm4s-memory and llm4s-memory-postgres
Second of the module carves tracked in
#1126; slice 2 is
#1129. It is in the build but not yet in a
release, so nothing here affects 0.4.1 or earlier.
What moved
| Packages | New module |
|---|---|
org.llm4s.agent.memory, except PostgresMemoryStore |
llm4s-memory |
org.llm4s.agent.memory.PostgresMemoryStore |
llm4s-memory-postgres |
Package names did not change, and there are no source breaks in this slice. Nothing
outside org.llm4s.agent.memory referenced it, so the whole package moved with no facade left
behind. Add the dependency; your imports stay as they are.
1
2
3
4
5
6
7
8
9
10
11
// Before
libraryDependencies += "org.llm4s" %% "llm4s-core" % version
// After — only if you use agent memory
libraryDependencies ++= Seq(
"org.llm4s" %% "llm4s-core" % version,
"org.llm4s" %% "llm4s-memory" % version
)
// ...and only if you store memories in Postgres/pgvector
libraryDependencies += "org.llm4s" %% "llm4s-memory-postgres" % version
Why two artifacts
PostgresMemoryStore was the only file in the package that needed a connection pool and a
server-side driver. Shipping it alongside InMemoryStore would mean every user of agent
memory inherits HikariCP and the Postgres JDBC driver whether or not they ever open a
connection. llm4s-memory carries sqlite-jdbc — for the file-backed SQLiteMemoryStore and
VectorMemoryStore — and nothing else; llm4s-memory-postgres depends on llm4s-memory and
adds the two heavy dependencies.
What llm4s-core sheds
HikariCP and the Postgres JDBC driver leave the core classpath. sqlite-jdbc leaves too:
all three used to be declared in the build’s shared settings, which put them on every
module’s classpath, so core could not shed them by itself. They are now declared only by the
modules that open a connection. If you depend on llm4s-core and use any of the three
directly, declare them yourself rather than relying on the transitive edge.
org.llm4s.vectorstore.PostgresVectorHelpers
Unchanged for callers — same package, same object, same methods — but worth knowing where it
ships. It is the pgvector text codec ([0.1,0.2,0.3] ⇄ Array[Float]), it names no JDBC
type, and it now has consumers in two modules that do not and should not depend on each
other: PgVectorStore in llm4s-rag and PostgresMemoryStore in llm4s-memory-postgres.
Rather than have either reach for the other, the single copy lives in llm4s-core, which both
already depend on. Slice 1 had briefly moved it into llm4s-rag and left a private duplicate
in core for PostgresMemoryStore; that duplicate is now gone.
So org.llm4s.vectorstore is split across two jars: this one object in llm4s-core, the rest
in llm4s-rag. It resolves the same way on any ordinary classpath.
Configuration keys
None. agent/memory reads no reference.conf keys and no Llm4sConfig method returns a type
that moved, so there is nothing to migrate.
Slice 1: llm4s-rag and llm4s-knowledgegraph
First of the module carves tracked in
#1126; slice 1 is
#1128. It is in the build but not yet in a
release, so nothing here affects 0.4.1 or earlier.
What moved
| Packages | New module |
|---|---|
org.llm4s.rag, org.llm4s.vectorstore, org.llm4s.chunking, org.llm4s.reranker, org.llm4s.eval, org.llm4s.extract, org.llm4s.knowledgegraph.graphrag |
llm4s-rag |
org.llm4s.knowledgegraph (everything except graphrag) |
llm4s-knowledgegraph |
Package names did not change, with the one deliberate exception described under Source breaks below. Add the dependency; your imports stay as they are.
1
2
3
4
5
6
7
8
// Before
libraryDependencies += "org.llm4s" %% "llm4s-core" % version
// After — only if you use RAG, vector stores, chunking, reranking or the knowledge graph
libraryDependencies ++= Seq(
"org.llm4s" %% "llm4s-core" % version,
"org.llm4s" %% "llm4s-rag" % version // depends on llm4s-knowledgegraph transitively
)
llm4s-rag depends on llm4s-knowledgegraph, so depending on the graph alone is only
worth doing if you want the graph without RAG.
llm4s-knowledgegraph-neo4j — a separate, already-published artifact — now depends on
llm4s-knowledgegraph instead of llm4s-core, which resolves for you.
What llm4s-core sheds
Six dependencies leave the core classpath: Tika, POI, PDFBox, jsoup, AWS S3 and AWS STS.
If you depend on llm4s-core and use any of those directly, declare them yourself rather
than relying on the transitive edge.
Source breaks
Three, all of them in this slice on purpose — pre-1.0 is when a duplicate is cheapest to remove.
1. The two document extractors are now one. org.llm4s.rag.extract.DocumentExtractor
and org.llm4s.llmconnect.extractors.UniversalExtractor were independent implementations of
one job: two Tika instances, two sets of MIME constants, two PDFBox paths, two POI paths.
They are now org.llm4s.extract.
| Before | After |
|---|---|
org.llm4s.rag.extract.DocumentExtractor |
org.llm4s.extract.DocumentExtractor |
org.llm4s.rag.extract.DefaultDocumentExtractor |
org.llm4s.extract.TikaDocumentExtractor |
UniversalExtractor.extract(path) → Either[ExtractorError, String] |
TikaDocumentExtractor.extractFromPath(path) → Result[ExtractedDocument] (text in .text) |
UniversalExtractor.extractFromBytes(bytes, name, mime) |
TikaDocumentExtractor.extract(bytes, name, mime) |
UniversalExtractor.extractFromStream(in, name, mime) |
TikaDocumentExtractor.extractFromStream(in, name, mime) (returns ExtractedDocument) |
UniversalExtractor.isTextLike(mime) |
TikaDocumentExtractor.canExtract(mime) — also true for legacy .doc |
UniversalExtractor.detectMimeType(bytes, name) |
unchanged |
UniversalExtractor.extractAny(path) and its Extracted / TextContent / ImageContent / AudioContent / VideoContent ADT |
org.llm4s.extract.MediaExtractor |
org.llm4s.llmconnect.model.ExtractorError |
org.llm4s.error.ProcessingError |
The package is org.llm4s.extract, not org.llm4s.rag.extract: extraction has two real
consumers — RAG document loading and multimodal embedding — and it quarantines the three
heaviest dependencies in the build. Naming it outside the rag namespace makes any later
decision to give it its own artifact a build-file change rather than a code change.
2. EmbeddingClient.encodePath is now FileEmbedder.encodeFromPath. EmbeddingClient
keeps the pure vector API; file reading, MIME sniffing and chunking live in
org.llm4s.rag.embed. The six-parameter signature — which included an
experimentalStubsEnabled: Boolean, a deployment decision arriving at a call site — became
a config object.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
// Before
client.encodePath(path, textModel, chunkingCfg, stubsEnabled, localModels)
// After
import org.llm4s.rag.embed.{ FileEmbedder, FileEmbeddingConfig, TextChunkingConfig }
FileEmbedder.encodeFromPath(
path,
client,
FileEmbeddingConfig(
textModel = textModel,
localModels = localModels,
chunking = TextChunkingConfig(enabled = true, size = 1000, overlap = 100),
experimentalStubs = stubsEnabled
)
)
UniversalEncoder.TextChunkingConfig is now the top-level org.llm4s.rag.embed.TextChunkingConfig.
3. Llm4sConfig.pgSearchIndex() is now PgSearchIndexConfigLoader.default(). It
returned a SearchIndex.PgConfig, which is RAG’s type; Llm4sConfig stays in llm4s-core
and cannot name it. The loader itself keeps its package (org.llm4s.config) and its
load(source) method, and moves to llm4s-rag.
1
2
3
4
5
6
// Before
val pg = Llm4sConfig.pgSearchIndex()
// After
import org.llm4s.config.PgSearchIndexConfigLoader
val pg = PgSearchIndexConfigLoader.default()
Configuration keys
llm4s.rag.permissions.pg.* and llm4s.rerank.* now ship in llm4s-rag’s reference.conf
rather than core’s. HOCON merges reference.conf across jars, so the key paths are
unchanged and nothing in your application.conf needs editing — but a build that reads
those keys without depending on llm4s-rag no longer gets the defaults.
llm4s.embeddings.* — including chunking and experimentalStubs — stays in core, because
Llm4sConfig still reads it there.
Artifact coordinate rename (v0.4.0)
Breaking change
Every published artifact under the org.llm4s group was renamed to carry an llm4s-
prefix and a consistent kebab-case suffix. This is a coordinate-only change: there are
no API changes, no package moves and no source changes in this release. Update your
build.sbt, recompile, and you are done.
| Old coordinate | New coordinate |
|---|---|
"org.llm4s" %% "core" |
"org.llm4s" %% "llm4s-core" |
"org.llm4s" %% "workspaceShared" (published as workspaceshared) |
"org.llm4s" %% "llm4s-workspace-shared" |
"org.llm4s" %% "workspaceClient" (published as workspaceclient) |
"org.llm4s" %% "llm4s-workspace-client" |
"org.llm4s" %% "trace-opentelemetry" |
"org.llm4s" %% "llm4s-observability-otel" |
"org.llm4s" %% "knowledgegraph-neo4j" |
"org.llm4s" %% "llm4s-knowledgegraph-neo4j" |
Maven users: the artifactId gains the same prefix, so core_3 becomes llm4s-core_3
(and core_2.13 becomes llm4s-core_2.13).
Why
Two reasons:
- Consistency.
org.llm4s:coreis a poor coordinate to read in somebody else’s build file, and the module names were an inconsistent mix of camelCase (silently lowercased by the publish intoworkspaceclient) and kebab-case. - Escaping a bad publish. The
core_3/core_2.13artifacts carry an accidental mis-published version2.1.593(a typo). Maven Central publishes are immutable, so that version cannot be retracted, and some resolvers sort it as the “latest” release. A fresh artifact name is the only way out; documentation is not.
Nobody is stranded
Releases up to and including 0.3.4 remain published, unchanged and resolvable under the
old coordinates. Pinning "org.llm4s" %% "core" % "0.3.4" keeps working indefinitely — you
only need to change coordinates when you move to 0.4.0 or later.
If you are pinning core with a floating or range version, pin an explicit 0.3.4 before
upgrading, so the phantom 2.1.593 is never selected.
Migration steps
- Replace the old coordinate with the new one in your
build.sbt(see the table above). - Set the version to
0.4.0or later. - Recompile. No imports, types or method signatures changed.
1
2
3
4
5
// Before
libraryDependencies += "org.llm4s" %% "core" % "0.3.4"
// After
libraryDependencies += "org.llm4s" %% "llm4s-core" % "0.4.0"
Note on "org.llm4s" %% "llm4s"
The aggregate llm4s artifact (llm4s_3 / llm4s_2.13) is not published and has not
been since 0.2.9 — the root project sets publish / skip := true. Any build file or
documentation that depends on "org.llm4s" %% "llm4s" is wrong independently of this
rename and should be changed to "org.llm4s" %% "llm4s-core".
Note on llm4s-observability-otel
The OpenTelemetry integration is published as llm4s-observability-otel rather than
llm4s-trace-opentelemetry. The name anticipates the llm4s-observability module that the
modularisation work will carve out of trace + metrics, so the integration is named once
rather than twice.
Unpublished modules
samples, workspaceRunner, workspaceSamples, config-policy, it and benchmarks set
publish / skip := true and never reached Maven Central. Their name values were made
consistent in the same change, but this has no effect on any downstream build.
MessageRole Enum Changes (v0.2.0)
Breaking Change
The MessageRole has been converted from string-based constants to a proper enum type for better type safety.
Before (v0.1.x)
1
2
3
4
5
6
7
8
import org.llm4s.llmconnect.model.Message
val message = Message(role = "assistant", content = "Hello")
message.role match {
case "assistant" => // handle assistant
case "user" => // handle user
case _ => // handle other
}
After (v0.2.0)
1
2
3
4
5
6
7
8
9
10
11
12
import org.llm4s.llmconnect.model.{Message, MessageRole}
val message = AssistantMessage(content = "Hello")
// or
val message = Message(role = MessageRole.Assistant, content = "Hello")
message.role match {
case MessageRole.Assistant => // handle assistant
case MessageRole.User => // handle user
case MessageRole.System => // handle system
case MessageRole.Tool => // handle tool
}
Migration Steps
- Update imports: Add
MessageRoleto your imports1
import org.llm4s.llmconnect.model.MessageRole
- Replace string comparisons: Update pattern matches and comparisons
1 2 3 4 5
// Before if (message.role == "assistant") { ... } // After if (message.role == MessageRole.Assistant) { ... }
- Update message creation: Use the typed constructors
1 2 3 4 5 6 7
// Before Message(role = "user", content = "Hello") // After UserMessage(content = "Hello") // or Message(role = MessageRole.User, content = "Hello")
Error Hierarchy Changes (v0.2.0)
New Error Categorization
Errors are now categorized using traits for better type safety and recovery strategies.
Before (v0.1.x)
1
2
3
4
error match {
case e: LLMError if e.isRecoverable => // retry logic
case e: LLMError => // handle non-recoverable
}
After (v0.2.0)
1
2
3
4
error match {
case e: RecoverableError => // retry logic
case e: NonRecoverableError => // handle non-recoverable
}
Error Recovery Pattern
1
2
3
4
5
6
7
8
9
10
import org.llm4s.error._
def handleError(error: LLMError): Unit = error match {
case _: RateLimitError => // wait and retry
case _: TimeoutError => // retry with backoff
case _: ServiceError with RecoverableError => // retry
case _: AuthenticationError => // refresh token or fail
case _: ValidationError => // fix input and retry
case _ => // non-recoverable, fail
}
Migration Steps
- Replace
isRecoverablechecks: Use pattern matching on traits1 2 3 4 5 6 7 8
// Before if (error.isRecoverable) { ... } // After error match { case _: RecoverableError => { ... } case _ => { ... } }
- Update error handling: Use the new trait-based categorization
1 2 3 4 5
// Before case e: ServiceError if e.isRecoverable => // After case e: ServiceError with RecoverableError =>
- Use smart constructors: Create errors using the companion object methods
1 2 3 4 5
// Before new RateLimitError(429, "Rate limit exceeded", Some(60.seconds)) // After RateLimitError(429, "Rate limit exceeded", Some(60.seconds))
Configuration Changes (v0.2.0+)
Llm4sConfig.provider()with no argument, used below, has since becomeLlm4sConfig.defaultProvider(), and providers are configured as named sections - see FromLLM_MODELto named provider sections.
EnvLoader and legacy ConfigReader → Llm4sConfig
Older versions used EnvLoader and a custom ConfigReader abstraction. These have been superseded by Llm4sConfig (PureConfig‑based) and typed helpers.
Before (v0.1.x)
1
2
3
4
import org.llm4s.config.EnvLoader
val apiKey = EnvLoader.get("OPENAI_API_KEY")
val model = EnvLoader.getOrElse("LLM_MODEL", "gpt-4")
or:
1
2
3
4
5
import org.llm4s.config.ConfigReader
import org.llm4s.llmconnect.LLMConnect
val client: org.llm4s.types.Result[org.llm4s.llmconnect.LLMClient] =
ConfigReader.Provider().flatMap(LLMConnect.getClient)
After (post‑0.2.0)
1
2
3
4
5
6
7
8
import org.llm4s.config.Llm4sConfig
import org.llm4s.llmconnect.LLMConnect
val client: org.llm4s.types.Result[org.llm4s.llmconnect.LLMClient] =
for {
cfg <- Llm4sConfig.provider()
client <- LLMConnect.getClient(cfg)
} yield client
Typed Config: recommended patterns
- Tracing (typed):
1 2 3 4 5
import org.llm4s.config.Llm4sConfig import org.llm4s.trace.{ Tracing, EnhancedTracing, TracingMode } val tracerResult: org.llm4s.types.Result[Tracing] = Llm4sConfig.tracing().map(Tracing.create)
- Provider model for display (typed):
1 2
val modelNameResult = Llm4sConfig.provider().map(_.model) // Prefer completion.model after the API call when available
- Workspace (samples):
1 2 3 4 5
import org.llm4s.codegen.WorkspaceConfigSupport val ws = WorkspaceConfigSupport.load().getOrElse( throw new IllegalArgumentException("Failed to load workspace settings") )
- Embeddings (samples):
1 2 3 4 5 6
val ui = org.llm4s.samples.embeddingsupport.EmbeddingUiSettings.loadFromEnv() .getOrElse(throw new IllegalArgumentException("Failed to load UI settings")) val targets = org.llm4s.samples.embeddingsupport.EmbeddingTargets.loadFromEnv() .fold(err => throw new IllegalArgumentException(err.toString), _.targets) val query = org.llm4s.samples.embeddingsupport.EmbeddingQuery.loadFromEnv() .fold(_ => None, _.value)
Configuration: legacy reader → Llm4sConfig / typed helpers (post‑0.2.0)
Earlier versions used a custom ConfigReader-style abstraction as a catch‑all for configuration. With PureConfig in place and typed helpers available, the preferred path is now:
- Use
org.llm4s.config.Llm4sConfigin core code. - Use explicit typed loaders plus
LLMConnect.getClientin application/sample code.
Provider configuration and client creation
Before (legacy reader-based API)
1
2
3
4
5
import org.llm4s.config.ConfigReader
import org.llm4s.llmconnect.LLMConnect
val client: org.llm4s.types.Result[org.llm4s.llmconnect.LLMClient] =
ConfigReader.Provider().flatMap(LLMConnect.getClient)
After
1
2
3
4
5
6
7
8
9
import org.llm4s.config.Llm4sConfig
import org.llm4s.llmconnect.LLMConnect
// Typed path using Llm4sConfig
val client: org.llm4s.types.Result[org.llm4s.llmconnect.LLMClient] =
for {
cfg <- Llm4sConfig.provider()
client <- LLMConnect.getClient(cfg)
} yield client
Tracing configuration
Before (legacy reader-based API)
1
2
3
4
5
import org.llm4s.config.ConfigReader
import org.llm4s.trace.Tracing
val tracer: Tracing =
ConfigReader.TracingConf().map(Tracing.create).getOrElse(Tracing.noop)
After
1
2
3
4
5
import org.llm4s.config.Llm4sConfig
import org.llm4s.trace.Tracing
val tracer: org.llm4s.types.Result[Tracing] =
Llm4sConfig.tracing().map(Tracing.create)
Embeddings: provider and client
Before (legacy reader-based API)
1
2
3
4
5
6
7
import org.llm4s.config.ConfigReader
import org.llm4s.llmconnect.EmbeddingClient
val client: org.llm4s.types.Result[EmbeddingClient] =
ConfigReader.Embeddings().flatMap { case (provider, cfg) =>
EmbeddingClient.from(provider, cfg)
}
After
1
2
3
4
5
6
7
import org.llm4s.config.Llm4sConfig
import org.llm4s.llmconnect.EmbeddingClient
val client: org.llm4s.types.Result[EmbeddingClient] =
Llm4sConfig.embeddings().flatMap { case (provider, cfg) =>
EmbeddingClient.from(provider, cfg)
}
Workspace settings
Before
1
2
3
4
5
import org.llm4s.codegen.WorkspaceSettings
val ws = WorkspaceSettings.load().getOrElse(
throw new IllegalArgumentException("Failed to load workspace settings")
)
After
1
2
3
4
5
import org.llm4s.codegen.WorkspaceConfigSupport
val ws = WorkspaceConfigSupport.load().getOrElse(
throw new IllegalArgumentException("Failed to load workspace settings")
)
API keys and types
Before (legacy reader-based API)
1
2
3
// Legacy pattern: API key resolved from a generic config reader
def loadApiKey(reader: /* legacy ConfigReader */ Any): Result[ApiKey] =
ApiKey.unsafe("sk-legacy-key") // placeholder for old behavior
After
1
2
3
import org.llm4s.config.Llm4sConfig
val cfgResult = Llm4sConfig.provider() // Result[ProviderConfig]
- For new code, do not introduce new parameters of reader/ConfigReader types. Prefer:
Llm4sConfigin core libraries.- Typed helpers plus
LLMConnect.getClient(andLlm4sConfig.tracing().map(Tracing.create)/.map(EnhancedTracing.create)for tracing) in applications and samples.
- For existing code that currently depends on a
ConfigReader-style abstraction:- Start by swapping call sites to use typed helpers (e.g.,
Llm4sConfig.provider()). - Where you need fine-grained control, switch to
Llm4sConfigfunctions instead of calling the legacy reader directly.
- Start by swapping call sites to use typed helpers (e.g.,