Agent Framework
Build sophisticated AI agents with tools, guardrails, memory, and multi-agent coordination.
Table of contents
- Overview
- Quick Start
- Safety Defaults
- Core Concepts
- Features
- Built-in Tools
- Context Window Management
- Conversation Persistence
- Reasoning Modes
- Examples
- Design Documents
- Next Steps
Overview
The LLM4S Agent Framework provides a production-ready foundation for building LLM-powered agents with:
- Tool Calling - Type-safe tools with automatic schema generation
- Guardrails - Input/output validation for safety and quality
- Memory - Short and long-term context with semantic search
- Handoffs - Agent-to-agent delegation for specialist routing
- Streaming - Real-time events for responsive UIs (
agent.stream, streaming guide) - Graphs - Typed multi-agent workflows on
GraphBuilder: parallel nodes, joins and checkpoints (recipe)
Quick Start
The agent runtime is the llm4s-agent module, alongside llm4s-core and a provider module:
1
libraryDependencies += "org.llm4s" %% "llm4s-agent" % llm4sVersion
Basic Agent
An agent is built once - tools, guardrails, handoffs and middleware belong to it - and run by thread.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
import org.llm4s.config.Llm4sConfig
import org.llm4s.llmconnect.LLMConnect
import org.llm4s.agent.Agent
// Create an agent and run a query
val result = for {
providerConfig <- Llm4sConfig.provider()
client <- LLMConnect.getClient(providerConfig)
agent <- Agent.builder("assistant", client).build()
result <- agent.run("What is the capital of France?")
} yield result
result match {
case Right(r) => println(r.answer.getOrElse(s"Run ended: ${r.status}"))
case Left(error) => println(s"Error: $error")
}
Agent.builder(id, client) takes the agent’s id (letters, digits, _ and -, up to 52) and an
LLMClient. build() returns a Result[Agent] and refuses an invalid agent: clashing tool names,
a bad or duplicate handoff id, or withMaxSteps below 1.
Agent with Tools
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
import org.llm4s.toolapi.{ToolRegistry, ToolFunction}
// Define a tool
def getWeather(location: String): String = {
s"The weather in $location is sunny, 72F"
}
val weatherTool = ToolFunction(
name = "get_weather",
description = "Get current weather for a location",
function = getWeather _
)
// Give the agent its tools when you build it
val result = for {
providerConfig <- Llm4sConfig.provider()
client <- LLMConnect.getClient(providerConfig)
agent <- Agent.builder("assistant", client)
.withTools(new ToolRegistry(Seq(weatherTool)))
.build()
result <- agent.run("What's the weather in Paris?")
} yield result
withTools takes a ToolRegistry (an MCPToolRegistry too) or a ToolSet of AgentTools.
Multi-Turn Conversations
A conversation is carried by its thread. continueConversation runs the next turn on the thread
of a previous result:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
val result = for {
providerConfig <- Llm4sConfig.provider()
client <- LLMConnect.getClient(providerConfig)
agent <- Agent.builder("assistant", client).build()
// First turn
result1 <- agent.run("Tell me about Scala")
// Follow-up (preserves context)
result2 <- agent.continueConversation(result1, "How does it compare to Java?")
// Another follow-up
result3 <- agent.continueConversation(result2, "What about performance?")
} yield result3
To name the thread yourself, use agent.run(threadId, query); agent.runMultiTurn(first, followUps)
runs several turns on one thread and stops at the first result that is not Completed. A Left carries no
thread id, so when a turn of run(query) or runMultiTurn fails or is cancelled, its thread is forgotten once
the turn has ended; name the thread to recover such a turn.
Handling the Result
run returns Result[AgentResult]. Left is a failure: a provider, tool or middleware error is a
GraphError, and a thread that is busy or still has pending work is refused. A run that ended in a
defined way is Right, with a status:
1
2
3
4
5
6
7
8
result.map { r =>
r.status match {
case AgentStatus.Completed(answer) => println(answer)
case AgentStatus.Blocked(guardrail, why) => println(s"Blocked by $guardrail: $why")
case AgentStatus.StepLimitReached => println("Hit the step limit")
case AgentStatus.Suspended(approvals, qs) => println("Waiting for a person")
}
}
AgentResult also carries threadId, runId, activeAgent, messages (the thread’s full
history, never a system prompt) and usage.
Safety Defaults
- Agent step limit: an agent makes at most
maxStepsmodel calls per turn,Agent.DefaultMaxSteps(50) unless you callwithMaxSteps(n). The count is shared by every agent a turn hands off to; at the limit the result isAgentStatus.StepLimitReached. The limit means the same throughrun,start,stream,continueConversationandrunMultiTurn: each turn gets its ownmaxStepsmodel calls. - HTTPTool methods:
HttpConfig()defaults toGETandHEADonly. UseHttpConfig.withWriteMethods()orHttpConfig().withAllMethodsto allow write methods.
Core Concepts
Threads and Results
A conversation lives in a thread on a GraphRuntime (in memory unless you call withRuntime). The
thread holds the messages, the active agent and the usage; it holds no tools, guardrails or agents,
which belong to the Agent. AgentResult is a value to read, not something to pass back in.
agent.start(threadId, query)returns anAgentRunat once, withawait()andcancel();runisstartfollowed byawait.- A run that failed or was cancelled leaves its thread recoverable:
agent.recover(threadId)re-runs only the work that did not finish. - A run that suspended for approval (
AgentStatus.Suspended) continues withagent.resume(threadId, answers), building each answer withresult.approve(id),result.reject(id, reason),result.edit(id, arguments)orresult.reply(id, value). - A thread stays in the runtime until
agent.forget(threadId)removes it - one-shotagent.run(query)threads too, which on the default in-memory runtime means they stay in memory. Forget a conversation you will not continue:agent.run(query).flatMap(r => agent.forget(r.threadId).map(_ => r)).forgetrefuses a thread whose run is still active (ThreadBusy) or that belongs to another tenant (TenantMismatch); an unknown thread isRight(()).
Suspended turns from Java and Kotlin
The Java facade (llm4s-java-api) returns every turn as a JAgentResult, read without Scala types:
answer() is an Optional<String>, messages() a java.util.List<JMessage>, usage() a
JUsageSummary, and status() a JAgentStatus whose kind() is the Java enum AgentStatusKind -
COMPLETED, BLOCKED, STEP_LIMIT_REACHED or SUSPENDED - with answer(), guardrail() and
reason() as Optional<String>s (see the Java guide). A SUSPENDED
status’s pending() is a java.util.List<PendingInterrupt>: the turn’s approvals, then its
questions; it is empty for any other status, and JAgent.pending(result) is a shortcut for it. Each PendingInterrupt has
id(), kind() (the Java enum InterruptKind, APPROVAL or QUESTION), toolName() and
argumentsJson(). An approval also has reason(), and a question has questionJson(), the tool’s
question as JSON. Both are Optional<String>, empty for the other kind. Answer each pending item
with Answer.approve(id), reject(id, reason), edit(id, argumentsJson) or reply(id, json), then
call agent.resume(threadId, answers), which blocks like run and returns an LlmResult<JAgentResult>.
You can answer only some of them: the result is SUSPENDED again, with the rest still pending.
1
2
3
4
5
6
7
8
9
10
11
12
LlmResult<JAgentResult> turn = agent.run("Deploy the release");
while (turn.get().status().kind() == AgentStatusKind.SUSPENDED) {
List<Answer> answers = new ArrayList<>();
for (PendingInterrupt p : turn.get().status().pending()) {
switch (p.kind()) {
case APPROVAL -> answers.add(Answer.approve(p.id())); // or reject / edit
case QUESTION -> answers.add(Answer.reply(p.id(), "{\"ok\":true}"));
}
}
turn = agent.resume(turn.get().threadId(), answers);
}
System.out.println(turn.get().answer().orElse("(" + turn.get().status().kind() + ")"));
agent.recover(threadId) continues a turn that failed or was cancelled, and also returns an
LlmResult<JAgentResult>. A failed result means one of these: a malformed answer, an empty answers
list (InvalidResume), an answer to an id the thread is not waiting for, a thread that is not suspended (resume), or a thread with nothing to
recover. resume and recover handle an interrupt the way run and continueConversation do: interrupting the
thread blocked in the call cancels the turn, and the call returns a CancelledError with the interrupt flag still
set once the turn has ended, leaving the thread for recover (see
Java Threading and Cancellation).
To cancel a turn without interrupting a thread, use streamResume or streamRecover and cancel the stream (see
Streaming Events).
The Kotlin API reuses these types: every AgentKt turn returns a JAgentResult, and when over
status().kind() covers it. AgentKt.pending(result) is the same List<PendingInterrupt>.
agent.resume(threadId, answers) and agent.recover(threadId) are suspend functions that return
the JAgentResult or throw LLMException. Every AgentKt suspend function that runs a turn - run,
continueConversation, resume and recover - runs it as the matching flow does (stream,
streamResume, streamRecover), so cancelling the caller - a cancelled scope, withTimeout -
cancels the turn. The call throws CancellationException once the turn has ended, and the
conversation thread is no longer busy: it is left for recover, which finishes the cancelled turn.
If the turn had already completed when the cancellation arrived, the call still throws
CancellationException, but the turn’s result is committed to the thread: recover then has nothing to
recover and throws LLMException (“no incomplete execution”), and the next turn continues from that result.
The one-shot run(query) is the exception, as in Java and Scala: nothing it throws carries its random thread
id, so once a failed or cancelled turn has ended it forgets the thread. A cancelled call waits up to 5 seconds
for the turn to end; a provider that ignores the interrupt for longer is left to finish on its own, and the
thread stays busy (ThreadBusy) until it does.
1
2
3
4
5
6
7
8
9
10
var turn = agent.run("Deploy the release")
while (turn.status().kind() == AgentStatusKind.SUSPENDED) {
val answers = turn.status().pending().map { p ->
when (p.kind()) {
InterruptKind.APPROVAL -> Answer.approve(p.id())
InterruptKind.QUESTION -> Answer.reply(p.id(), """{"ok":true}""")
}
}
turn = agent.resume(turn.threadId(), answers)
}
Agent Lifecycle
1
2
3
4
5
6
7
8
9
10
11
12
13
Query
|
v
input node (input guardrails) ---> Blocked (nothing stored)
|
v
model <---------- tool results
|
+-- tool calls --> call-tool (parallel, bounded) --> collect --> model
+-- handoff -----> the target agent's model
|
v
finish node (output guardrails) ---> Completed | Blocked (turn removed)
Parallel Tool Calls
The tool calls of one model message run in parallel, bounded by RunBudgets.maxConcurrency,
which you set through the RunConfig you pass to run.
ToolExecutionStrategy remains in core as a ToolRegistry feature; the agent no longer takes one.
Features
Guardrails
Validate inputs and outputs for safety. Guardrails are middleware on the agent:
1
2
3
4
5
6
7
8
9
10
11
import org.llm4s.agent.graph.middleware.GuardrailMiddleware
import org.llm4s.agent.guardrails.builtin._
val agent = Agent.builder("assistant", client)
.withMiddleware(
new GuardrailMiddleware(
input = Seq(new LengthCheck(1, 10000), new ProfanityFilter()),
output = Seq(new JSONValidator())
)
)
.build()
A blocked run ends AgentStatus.Blocked(guardrail, reason), the blocked turn removed from the thread, which stays usable.
Memory System
Persistent context across conversations:
1
2
3
4
5
6
7
import org.llm4s.agent.memory._
val result = for {
manager <- SimpleMemoryManager.empty
m1 <- manager.recordUserFact("Prefers Scala", Some("user-1"), Some(0.9))
context <- m1.getRelevantContext("programming preferences")
} yield context
Handoffs
Delegate to specialist agents:
1
2
3
4
5
6
7
8
import org.llm4s.agent.Handoff
val physics = Agent.builder("physics", client)
.withSystemPrompt("You are a physicist")
val agent = Agent.builder("triage", client)
.withHandoffs(Handoff.to("physics", physics, "Physics expertise required"))
.build()
Streaming Events
Agent.builder(...).withStreaming() streams the model’s answer, and
agent.stream(threadId, query)(listener) delivers every event of the turn - text deltas, tool
calls, model calls, handoffs, guardrail blocks - as a StreamEvent, matched with AgentEvents.
Durable events carry no message content; content is live-only. AgentIO.stream and AgentZ.stream
give the same events as fs2 and ZIO streams.
Durable Checkpointers
Every run claims its thread in the Checkpointer before it starts, renews the claim while it runs, and fences
every commit with the claim’s token. Runtimes sharing one store - in one process or many - are refused a live
run’s thread with ThreadBusy, can recover a run whose process died once its claim has expired, and a run that
lost its claim records nothing more. CheckpointerContract in llm4s-agent-testkit tests a store of your own.
Learn more about durable checkpointers →
Built-in Tools
The llm4s-agent-tools module ships ready-made tools - a calculator, the date and time, UUIDs, JSON, files, HTTP, a
shell and web search - in bundles you register with an agent. They work with plain tool calling through
ToolRegistry too, without an Agent. The Built-in Tools guide lists every tool and its
parameters, and says what each bundle lets a model do: read its safety section before you give a model files, the
network or a shell. For the module itself, see
installation.
1
2
3
4
5
6
7
8
9
import org.llm4s.agent.Agent
import org.llm4s.toolapi.ToolRegistry
import org.llm4s.toolapi.builtin.BuiltinTools
val result = for {
tools <- BuiltinTools.coreSafe // date/time, calculator, UUID, JSON: no files, network or processes
agent <- Agent.builder("assistant", client).withTools(new ToolRegistry(tools)).build()
state <- agent.run("What is 15% of 850?")
} yield state
The bundles, smallest to largest: coreSafe, withHttpSafe() (adds read-only HTTP), withFilesSafe() (adds
read-only files) and developmentSafe() (adds file writing and a shell). customSafe(...) builds exactly the tools
whose configuration you pass.
Context Window Management
Handle long conversations with ContextWindowMiddleware. It prunes what is sent to the model on
each call; the thread keeps its full history:
1
2
3
4
5
6
7
8
9
10
11
12
13
import org.llm4s.agent.{ContextWindowConfig, PruningStrategy}
import org.llm4s.agent.graph.middleware.ContextWindowMiddleware
val config = ContextWindowConfig(
maxMessages = Some(20),
preserveSystemMessage = true,
minRecentTurns = 2,
pruningStrategy = PruningStrategy.OldestFirst
)
val agent = Agent.builder("assistant", client)
.withMiddleware(new ContextWindowMiddleware(config))
.build()
The current turn (the latest user message onward) is never pruned, a tool call is never separated from its result, and the request always starts with a user message - so a request may exceed the budget. The system prompt is outside the budget.
Pruning Strategies:
| Strategy | Behavior |
|---|---|
OldestFirst |
Remove oldest messages first (FIFO) |
MiddleOut |
Keep first and last messages, remove middle |
RecentTurnsOnly(n) |
Keep only the last N conversation turns |
Custom(fn) |
User-defined pruning function |
Conversation Persistence
A thread lives in the runtime. To keep a conversation across processes, save the result’s messages
and import them as the history of a new thread:
1
2
3
4
5
6
7
8
import upickle.default.{read, write}
// Save the messages of a completed run
Files.writeString(path, write(result.messages))
// Later: load them and continue on a new thread
val saved = read[Vector[Message]](Files.readString(path))
val next = agent.run(ThreadId(java.util.UUID.randomUUID().toString), "Continue our conversation", RunConfig(), saved)
history is accepted only on a thread that does not exist yet, and may not contain system messages
(prompts belong to the agents). Save only runs that completed; history that ends mid tool call is
refused.
Reasoning Modes
Enable extended thinking for complex problems:
1
2
3
4
5
6
7
8
import org.llm4s.llmconnect.model.{CompletionOptions, ReasoningEffort}
val options = CompletionOptions()
.withReasoning(ReasoningEffort.High) // None, Low, Medium, High
.withMaxTokens(4096)
// Use with agent
Agent.builder("assistant", client).withCompletionOptions(options).build()
Supported by OpenAI o1/o3 and Anthropic Claude models.
Examples
| Example | Description |
|---|---|
| SingleStepAgentExample | A run with one tool call |
| MultiStepAgentExample | Complete execution flow |
| MultiTurnConversationExample | Functional multi-turn API |
| LongConversationExample | Context window pruning |
| ConversationPersistenceExample | Save and resume |
| BuiltinToolsAgentExample | Built-in tools |
Design Documents
For in-depth technical details:
- Agent Framework Roadmap - Strategic direction
- Phase 1.1: Conversations - Conversation API
- Phase 1.2: Guardrails - Validation framework
- Phase 1.3: Handoffs - Agent delegation
- Phase 1.4: Memory - Memory architecture
- Typed Agent Runtime §4.13 - How
Agentruns on the graph runtime - Phase 2.1: Streaming - Event system
Next Steps
- Guardrails Guide - Input/output validation
- Memory Guide - Persistent context
- Handoffs Guide - Agent delegation
- Streaming Guide - Real-time events
- Durable Checkpointers - Shared stores, claims and recovery
- Examples Gallery - Working code samples