Using LLM4S from Java

Call a model from Java with llm4s-java-api: a client, a prompt, a conversation, and failures you can read without Scala.

Table of contents

  1. What llm4s-java-api is
  2. Add the dependency
  3. Configure a provider
    1. Configure in code
  4. Your first call
    1. The snippets’ imports
  5. A conversation
  6. Reading a result
  7. Handling a failure
  8. An agent turn
  9. What is not here yet
  10. Where next

What llm4s-java-api is

LLM4S is written in Scala. Its core API returns Scala types (Either, Option, Seq) that are awkward to use from Java. llm4s-java-api is a small layer in front of it for Java callers: Llm4s creates a client, JLlmClient calls the model, ConversationBuilder builds a conversation, and every call returns an LlmResult, a result type modelled on java.util.Optional and CompletableFuture. A failed call is a value you check, not an exception you have to catch.

It does not hide every Scala type. Conversation, CompletionOptions, ProviderConfig and LLMError are Scala classes that still appear in its signatures. What is not here yet says which of them gets in your way.

Not published yet. llm4s-java-api is not on Maven Central: the latest release, 0.4.1, is a single llm4s-core artifact, and the Java API is planned to ship with 0.5.0. Until then, build it from a checkout of the repository; the gradle-java sample has the steps (its “Run it” section) and is a complete Java project you can copy. Which JDK is the minimum is not settled either (CI runs JDK 21 only today): follow #1493.

Add the dependency

The artifact is org.llm4s:llm4s-java-api_3. The _3 is the Scala binary version, and LLM4S is Scala 3 only. sbt adds it for you (%%); Maven and Gradle do not, so write it yourself.

Maven

1
2
3
4
5
<dependency>
    <groupId>org.llm4s</groupId>
    <artifactId>llm4s-java-api_3</artifactId>
    <version>0.4.1</version>
</dependency>

Gradle (Kotlin DSL)

1
2
3
dependencies {
    implementation("org.llm4s:llm4s-java-api_3:0.4.1")
}

sbt

1
libraryDependencies += "org.llm4s" %% "llm4s-java-api" % "0.4.1"

What arrives with it:

  • llm4s-core, llm4s-agent and the OpenAI, Anthropic, Ollama, Gemini and OpenAI-compatible provider modules, so you do not add a provider artifact. Making providers opt-in is tracked in #1496.
  • Scala’s standard library (scala3-library_3 and scala-library). If your build pins or excludes it, see the Gradle guide for the recipes that keep the two Scala libraries aligned.
  • No logging backend. LLM4S logs through SLF4J; add one (for example Logback), or SLF4J prints a warning and drops the logs.

Configure a provider

Java code does not build a provider configuration. It comes from an application.conf on your classpath, the same file Scala users write: a named section per provider, and one section named as the default.

1
2
3
4
5
6
7
8
9
10
11
# src/main/resources/application.conf
llm4s {
  providers {
    provider = "openai-main"          # the default: the name of a section below

    openai-main {
      provider = "openai"
      model    = "gpt-4o-mini"
    }
  }
}

With OPENAI_API_KEY set in the environment, that is the whole configuration: the OpenAI module binds the variable to the shared credential for openai, so the section needs only provider and model. Other providers work the same way; anthropic reads ANTHROPIC_API_KEY, and ollama needs no key but does need a baseUrl:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
llm4s {
  providers {
    provider = "claude"

    claude {                              # key from ANTHROPIC_API_KEY
      provider = "anthropic"
      model    = "claude-sonnet-4-20250514"
    }

    ollama-local {
      provider = "ollama"
      model    = "llama3.2"
      baseUrl  = "http://localhost:11434"
    }
  }
}

The configuration guide lists every key and every provider. Nothing in the library reads a LLM_MODEL variable; the model is the model key of the section.

Configure in code

Llm4s.createClient(ProviderConfig) builds a client from a config object instead of a file. For OpenAI, Anthropic and Ollama the config classes have a Java entry point that takes the essentials and fills in the rest (the default base URL, and a context window and completion reserve guessed from the model name):

1
2
3
LlmResult<JLlmClient> openai = Llm4s.createClient(OpenAIConfig.apply(apiKey, "gpt-4o-mini"));
LlmResult<JLlmClient> anthropic = Llm4s.createClient(AnthropicConfig.apply(apiKey, "claude-sonnet-4-20250514"));
LlmResult<JLlmClient> local = Llm4s.createClient(OllamaConfig.apply("llama3.2", "http://localhost:11434"));

(The imports are listed under The snippets’ imports.) Reading the key is up to you; creating the client sends no request. The with methods of each config (withContextWindow, withReserveCompletion, and for OpenAI withBaseUrl and withOrganization) change one field at a time. The other providers (Gemini, Azure, the OpenAI-compatible family and so on) have no such entry point yet, so configure them in application.conf.

Your first call

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
import java.io.PrintStream;

import org.llm4s.javaapi.JLlmClient;
import org.llm4s.javaapi.Llm4s;
import org.llm4s.javaapi.LlmResult;

public final class HelloLlm4s {

    private HelloLlm4s() {}

    public static void main(String[] args) {
        System.exit(run(System.out, System.err));
    }

    static int run(PrintStream out, PrintStream err) {
        LlmResult<JLlmClient> created = Llm4s.createDefaultClient();
        if (created.isFailure()) {
            err.println("Could not create a client: " + created.getError().getMessage());
            return 1;
        }

        // JLlmClient is AutoCloseable: it holds an HTTP client, so close it.
        try (JLlmClient client = created.get()) {
            LlmResult<String> answer = client.complete("What is a monad? Answer in one sentence.");
            answer
                .ifSuccess(text -> out.println(text))
                .ifFailure(error -> err.println("The call failed: " + error.getMessage()));
            return answer.isSuccess() ? 0 : 1;
        }
    }
}

Llm4s.createDefaultClient() loads the default provider section and builds a client from it. It does not throw: a missing section, a missing API key or an unknown provider is a failed LlmResult, and its message says what to fix (for a missing key it names the environment variable and the section). JLlmClient is AutoCloseable and closing it closes the underlying client, so use try-with-resources.

ifSuccess and ifFailure each take a lambda and return the result, so they chain. The run method is separate from main so that a test can pass its own streams.

The gradle-java sample is this program in a complete Gradle project.

The snippets’ imports

The snippets below leave out their imports. Together they use these:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
import java.util.Optional;
import java.util.concurrent.CompletableFuture;

import org.llm4s.error.LLMError;
import org.llm4s.error.RecoverableError;
import org.llm4s.javaapi.ConversationBuilder;
import org.llm4s.javaapi.JLlmClient;
import org.llm4s.javaapi.Llm4s;
import org.llm4s.javaapi.LlmException;
import org.llm4s.javaapi.LlmResult;
import org.llm4s.llmconnect.config.AnthropicConfig;
import org.llm4s.llmconnect.config.OllamaConfig;
import org.llm4s.llmconnect.config.OpenAIConfig;
import org.llm4s.llmconnect.model.Conversation;

A conversation

A conversation is a list of messages. ConversationBuilder builds one without Scala syntax:

1
2
3
4
5
6
7
Conversation conversation = ConversationBuilder.create()
    .system("You answer in one short sentence.")
    .user("What is a monad?")
    .build();

LlmResult<String> answer = client.complete(conversation);
System.out.println(answer.get());

The builder has system, user and assistant, and messages are sent in the order you add them. It is immutable: each method returns a new builder, so one can be shared and extended. A null message throws NullPointerException at once.

Reading a result

Every call returns an LlmResult<String>. There are several ways to take the value out, depending on how you want to treat a failure:

1
2
3
4
5
6
7
LlmResult<String> result = client.complete("What is 2+2?");

String text = result.get();                       // the value, or throws LlmException
String orNull = result.getOrNull();               // the value, or null when the call failed
Optional<String> optional = result.toOptional();  // Optional.empty() when the call failed
LlmResult<Integer> length = result.map(String::length);
CompletableFuture<String> future = result.toCompletableFuture();

isSuccess() and isFailure() test it, getError() returns the LlmException of a failure (or null), and map transforms a success and passes a failure through. toCompletableFuture() returns a future that is already complete: it makes a result fit an API that expects a future, and does not make the call asynchronous (see What is not here yet).

Handling a failure

A call that fails does not throw. JLlmClient turns a failure of the provider, the network or the library into a failed LlmResult, and a null argument into a failed result too. Only get() throws, and it throws LlmException, an unchecked exception that carries the llm4s error:

1
2
3
4
5
6
7
8
9
10
11
try {
    String text = client.complete("What is 2+2?").get();
    System.out.println(text);
} catch (LlmException e) {
    LLMError error = e.error();                // the llm4s error: a Scala type
    System.err.println(error.message());       // the same text as e.getMessage()
    System.err.println(error.formatted());     // the message plus its code and context
    if (error instanceof RecoverableError) {
        System.err.println("a retry may succeed");
    }
}

LLMError is a Scala trait, but message() and formatted() are plain methods you can call from Java. If the error carries a Throwable, it is the exception’s getCause().

Errors are classes you can test with instanceof. The ones that may succeed if tried again, perhaps after you do something first, implement RecoverableError: RateLimitError, TimeoutError, NetworkError, APIError and ServiceError among them. AuthenticationError, ConfigurationError and ValidationError implement NonRecoverableError: retrying the same request will not help. The error handling guide has the full list and the recovery tools, written for Scala.

If the thread blocked in complete is interrupted, the call returns a failed result whose error is a CancelledError, with the thread’s interrupt flag still set. InterruptedException is never thrown, so the method does not declare it and Java will not let you write catch (InterruptedException e) around the call; test the result for a CancelledError instead. The threading and cancellation guide has the details.

An agent turn

Llm4s.createAgent(client) wraps a client in a JAgent. run and continueConversation block, like complete, and return an LlmResult<JAgentResult>. A JAgentResult is read with JDK types and this module’s own: answer() is an Optional<String>, messages() a java.util.List<JMessage>, and status().kind() the Java enum AgentStatusKind, so a switch covers every way a turn ends:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
JAgent agent = Llm4s.createAgent(client);
JAgentResult result = agent.run("What is 2+2?").get();

switch (result.status().kind()) {
    case COMPLETED -> System.out.println(result.answer().orElseThrow());
    case BLOCKED -> System.out.println("Blocked by " + result.status().guardrail().orElseThrow());
    case STEP_LIMIT_REACHED -> System.out.println("Hit the step limit");
    case SUSPENDED -> System.out.println("Waiting for " + result.status().pending().size() + " answers");
}
for (JMessage message : result.messages()) {
    System.out.println(message.role() + ": " + message.content());
}
JUsageSummary usage = result.usage();
System.out.println(usage.inputTokens() + " tokens in, " + usage.outputTokens() + " out");

JAgentResult next = agent.continueConversation(result, "And 3+3?").get();

Each status’s data has its own accessor, empty for the other kinds: answer() for COMPLETED, guardrail() and reason() for BLOCKED, and pending() for SUSPENDED - the approvals and questions the turn waits for, which the agent guide’s suspended turns section answers. A JMessage has a role() (the Java enum JMessageRole), its content(), the toolCalls() an assistant message asked for, with their arguments as JSON text (an object, as a model sends them; a call built with a ujson.Str renders as a JSON string literal), and the toolCallId() a tool message answers. usage() counts tokens as longs, the cost as a java.math.BigDecimal (two usages are equal when their costs are numerically equal, whatever the scale), and byModel() breaks both down per model. No accessor returns a Scala or ujson type.

These types are values, but not Serializable. Their toString prints the full text - a message’s content, a tool call’s arguments, an answer or a guardrail’s reason - as the Scala types do, so mind what you log.

What is not here yet

llm4s-java-api covers a client and a conversation. These are not available from Java yet, each with the issue that tracks it:

You may expect State today
Completion options (temperature, max tokens, reasoning) complete(Conversation, CompletionOptions) exists, but a CompletionOptions takes scala.Option and Seq arguments to construct: #1488
Streaming tokens JLlmClient has blocking complete calls only: #1485
An asynchronous call none; wrap complete yourself, or use the Spring Boot starter’s completeAsync. Threading and cancellation are being documented in #1500
Structured output into a Java record #1486
Defining tools an agent takes a Scala ToolRegistry: #1484
Agents beyond a turn An agent turn reads a result; the agent guide covers streaming a turn and suspended turns from Java. Tools still need a Scala ToolRegistry (row above)
Embeddings and RAG #1490, #1491
A fake client for your own tests #1497

Where next