Using LLM4S from Java
Call a model from Java with llm4s-java-api: a client, a prompt, a conversation, and failures you can read without Scala.
Table of contents
- What
llm4s-java-apiis - Add the dependency
- Configure a provider
- Your first call
- A conversation
- Reading a result
- Handling a failure
- An agent turn
- What is not here yet
- Where next
What llm4s-java-api is
LLM4S is written in Scala. Its core API returns Scala types (Either, Option, Seq) that are awkward to use from Java.
llm4s-java-api is a small layer in front of it for Java callers: Llm4s creates a client, JLlmClient calls the
model, ConversationBuilder builds a conversation, and every call returns an LlmResult, a result type modelled on
java.util.Optional and CompletableFuture. A failed call is a value you check, not an exception you have to catch.
It does not hide every Scala type. Conversation, CompletionOptions, ProviderConfig and LLMError are Scala
classes that still appear in its signatures. What is not here yet says which of them gets in
your way.
Not published yet.
llm4s-java-apiis not on Maven Central: the latest release,0.4.1, is a singlellm4s-coreartifact, and the Java API is planned to ship with0.5.0. Until then, build it from a checkout of the repository; thegradle-javasample has the steps (its “Run it” section) and is a complete Java project you can copy. Which JDK is the minimum is not settled either (CI runs JDK 21 only today): follow #1493.
Add the dependency
The artifact is org.llm4s:llm4s-java-api_3. The _3 is the Scala binary version, and LLM4S is Scala 3 only. sbt adds
it for you (%%); Maven and Gradle do not, so write it yourself.
Maven
1
2
3
4
5
<dependency>
<groupId>org.llm4s</groupId>
<artifactId>llm4s-java-api_3</artifactId>
<version>0.4.1</version>
</dependency>
Gradle (Kotlin DSL)
1
2
3
dependencies {
implementation("org.llm4s:llm4s-java-api_3:0.4.1")
}
sbt
1
libraryDependencies += "org.llm4s" %% "llm4s-java-api" % "0.4.1"
What arrives with it:
llm4s-core,llm4s-agentand the OpenAI, Anthropic, Ollama, Gemini and OpenAI-compatible provider modules, so you do not add a provider artifact. Making providers opt-in is tracked in #1496.- Scala’s standard library (
scala3-library_3andscala-library). If your build pins or excludes it, see the Gradle guide for the recipes that keep the two Scala libraries aligned. - No logging backend. LLM4S logs through SLF4J; add one (for example Logback), or SLF4J prints a warning and drops the logs.
Configure a provider
Java code does not build a provider configuration. It comes from an application.conf on your classpath, the same file
Scala users write: a named section per provider, and one section named as the default.
1
2
3
4
5
6
7
8
9
10
11
# src/main/resources/application.conf
llm4s {
providers {
provider = "openai-main" # the default: the name of a section below
openai-main {
provider = "openai"
model = "gpt-4o-mini"
}
}
}
With OPENAI_API_KEY set in the environment, that is the whole configuration: the OpenAI module binds the variable to
the shared credential for openai, so the section needs only provider and model. Other providers work the same way;
anthropic reads ANTHROPIC_API_KEY, and ollama needs no key but does need a baseUrl:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
llm4s {
providers {
provider = "claude"
claude { # key from ANTHROPIC_API_KEY
provider = "anthropic"
model = "claude-sonnet-4-20250514"
}
ollama-local {
provider = "ollama"
model = "llama3.2"
baseUrl = "http://localhost:11434"
}
}
}
The configuration guide lists every key and every provider. Nothing in the library
reads a LLM_MODEL variable; the model is the model key of the section.
Configure in code
Llm4s.createClient(ProviderConfig) builds a client from a config object instead of a file. For OpenAI, Anthropic and
Ollama the config classes have a Java entry point that takes the essentials and fills in the rest (the default base
URL, and a context window and completion reserve guessed from the model name):
1
2
3
LlmResult<JLlmClient> openai = Llm4s.createClient(OpenAIConfig.apply(apiKey, "gpt-4o-mini"));
LlmResult<JLlmClient> anthropic = Llm4s.createClient(AnthropicConfig.apply(apiKey, "claude-sonnet-4-20250514"));
LlmResult<JLlmClient> local = Llm4s.createClient(OllamaConfig.apply("llama3.2", "http://localhost:11434"));
(The imports are listed under The snippets’ imports.) Reading the key is up to you; creating the client sends no
request. The with methods of each config (withContextWindow, withReserveCompletion, and for OpenAI
withBaseUrl and withOrganization) change one field at a time. The other providers (Gemini, Azure, the
OpenAI-compatible family and so on) have no such entry point yet, so configure them in application.conf.
Your first call
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
import java.io.PrintStream;
import org.llm4s.javaapi.JLlmClient;
import org.llm4s.javaapi.Llm4s;
import org.llm4s.javaapi.LlmResult;
public final class HelloLlm4s {
private HelloLlm4s() {}
public static void main(String[] args) {
System.exit(run(System.out, System.err));
}
static int run(PrintStream out, PrintStream err) {
LlmResult<JLlmClient> created = Llm4s.createDefaultClient();
if (created.isFailure()) {
err.println("Could not create a client: " + created.getError().getMessage());
return 1;
}
// JLlmClient is AutoCloseable: it holds an HTTP client, so close it.
try (JLlmClient client = created.get()) {
LlmResult<String> answer = client.complete("What is a monad? Answer in one sentence.");
answer
.ifSuccess(text -> out.println(text))
.ifFailure(error -> err.println("The call failed: " + error.getMessage()));
return answer.isSuccess() ? 0 : 1;
}
}
}
Llm4s.createDefaultClient() loads the default provider section and builds a client from it. It does not throw: a
missing section, a missing API key or an unknown provider is a failed LlmResult, and its message says what to fix (for
a missing key it names the environment variable and the section). JLlmClient is AutoCloseable and closing it closes
the underlying client, so use try-with-resources.
ifSuccess and ifFailure each take a lambda and return the result, so they chain. The run method is separate from
main so that a test can pass its own streams.
The gradle-java sample is this program in a
complete Gradle project.
The snippets’ imports
The snippets below leave out their imports. Together they use these:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
import java.util.Optional;
import java.util.concurrent.CompletableFuture;
import org.llm4s.error.LLMError;
import org.llm4s.error.RecoverableError;
import org.llm4s.javaapi.ConversationBuilder;
import org.llm4s.javaapi.JLlmClient;
import org.llm4s.javaapi.Llm4s;
import org.llm4s.javaapi.LlmException;
import org.llm4s.javaapi.LlmResult;
import org.llm4s.llmconnect.config.AnthropicConfig;
import org.llm4s.llmconnect.config.OllamaConfig;
import org.llm4s.llmconnect.config.OpenAIConfig;
import org.llm4s.llmconnect.model.Conversation;
A conversation
A conversation is a list of messages. ConversationBuilder builds one without Scala syntax:
1
2
3
4
5
6
7
Conversation conversation = ConversationBuilder.create()
.system("You answer in one short sentence.")
.user("What is a monad?")
.build();
LlmResult<String> answer = client.complete(conversation);
System.out.println(answer.get());
The builder has system, user and assistant, and messages are sent in the order you add them. It is immutable: each
method returns a new builder, so one can be shared and extended. A null message throws NullPointerException at once.
Reading a result
Every call returns an LlmResult<String>. There are several ways to take the value out, depending on how you want to
treat a failure:
1
2
3
4
5
6
7
LlmResult<String> result = client.complete("What is 2+2?");
String text = result.get(); // the value, or throws LlmException
String orNull = result.getOrNull(); // the value, or null when the call failed
Optional<String> optional = result.toOptional(); // Optional.empty() when the call failed
LlmResult<Integer> length = result.map(String::length);
CompletableFuture<String> future = result.toCompletableFuture();
isSuccess() and isFailure() test it, getError() returns the LlmException of a failure (or null), and map
transforms a success and passes a failure through. toCompletableFuture() returns a future that is already complete: it
makes a result fit an API that expects a future, and does not make the call asynchronous (see
What is not here yet).
Handling a failure
A call that fails does not throw. JLlmClient turns a failure of the provider, the network or the library into a failed
LlmResult, and a null argument into a failed result too. Only get() throws, and it throws LlmException, an
unchecked exception that carries the llm4s error:
1
2
3
4
5
6
7
8
9
10
11
try {
String text = client.complete("What is 2+2?").get();
System.out.println(text);
} catch (LlmException e) {
LLMError error = e.error(); // the llm4s error: a Scala type
System.err.println(error.message()); // the same text as e.getMessage()
System.err.println(error.formatted()); // the message plus its code and context
if (error instanceof RecoverableError) {
System.err.println("a retry may succeed");
}
}
LLMError is a Scala trait, but message() and formatted() are plain methods you can call from Java. If the error
carries a Throwable, it is the exception’s getCause().
Errors are classes you can test with instanceof. The ones that may succeed if tried again, perhaps after you do
something first, implement RecoverableError: RateLimitError, TimeoutError, NetworkError, APIError and
ServiceError among them. AuthenticationError, ConfigurationError and ValidationError implement
NonRecoverableError: retrying the same request will not help. The error handling guide has the full
list and the recovery tools, written for Scala.
If the thread blocked in complete is interrupted, the call returns a failed result whose error is a CancelledError,
with the thread’s interrupt flag still set. InterruptedException is never thrown, so the method does not declare it
and Java will not let you write catch (InterruptedException e) around the call; test the result for a
CancelledError instead. The threading and cancellation guide has the details.
An agent turn
Llm4s.createAgent(client) wraps a client in a JAgent. run and continueConversation block, like complete, and
return an LlmResult<JAgentResult>. A JAgentResult is read with JDK types and this module’s own: answer() is an
Optional<String>, messages() a java.util.List<JMessage>, and status().kind() the Java enum AgentStatusKind, so
a switch covers every way a turn ends:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
JAgent agent = Llm4s.createAgent(client);
JAgentResult result = agent.run("What is 2+2?").get();
switch (result.status().kind()) {
case COMPLETED -> System.out.println(result.answer().orElseThrow());
case BLOCKED -> System.out.println("Blocked by " + result.status().guardrail().orElseThrow());
case STEP_LIMIT_REACHED -> System.out.println("Hit the step limit");
case SUSPENDED -> System.out.println("Waiting for " + result.status().pending().size() + " answers");
}
for (JMessage message : result.messages()) {
System.out.println(message.role() + ": " + message.content());
}
JUsageSummary usage = result.usage();
System.out.println(usage.inputTokens() + " tokens in, " + usage.outputTokens() + " out");
JAgentResult next = agent.continueConversation(result, "And 3+3?").get();
Each status’s data has its own accessor, empty for the other kinds: answer() for COMPLETED, guardrail() and
reason() for BLOCKED, and pending() for SUSPENDED - the approvals and questions the turn waits for, which the
agent guide’s suspended turns section answers. A JMessage has a
role() (the Java enum JMessageRole), its content(), the toolCalls() an assistant message asked for, with their
arguments as JSON text (an object, as a model sends them; a call built with a ujson.Str renders as a JSON string
literal), and the toolCallId() a tool message answers. usage() counts tokens as longs, the cost as a
java.math.BigDecimal (two usages are equal when their costs are numerically equal, whatever the scale), and
byModel() breaks both down per model. No accessor returns a Scala or ujson type.
These types are values, but not Serializable. Their toString prints the full text - a message’s content, a tool
call’s arguments, an answer or a guardrail’s reason - as the Scala types do, so mind what you log.
What is not here yet
llm4s-java-api covers a client and a conversation. These are not available from Java yet, each with the issue that
tracks it:
| You may expect | State today |
|---|---|
| Completion options (temperature, max tokens, reasoning) | complete(Conversation, CompletionOptions) exists, but a CompletionOptions takes scala.Option and Seq arguments to construct: #1488 |
| Streaming tokens | JLlmClient has blocking complete calls only: #1485 |
| An asynchronous call | none; wrap complete yourself, or use the Spring Boot starter’s completeAsync. Threading and cancellation are being documented in #1500 |
| Structured output into a Java record | #1486 |
| Defining tools | an agent takes a Scala ToolRegistry: #1484 |
| Agents beyond a turn | An agent turn reads a result; the agent guide covers streaming a turn and suspended turns from Java. Tools still need a Scala ToolRegistry (row above) |
| Embeddings and RAG | #1490, #1491 |
| A fake client for your own tests | #1497 |
Where next
- Spring Boot: auto-configuration, a template bean and a health indicator on top of this module.
- Basic usage and providers: the same concepts in Scala, and every provider’s configuration.
- Gradle integration: dependency recipes for Gradle builds.
- The
gradle-javasample: a Java project you can run.