Jev, TypeSafe’s decision model
llm4s-jev is a client for Jev, TypeSafe’s System One model. Jev is not a chat
model. You send it a state and named, typed questions, and it answers each question with a typed answer your code can
act on directly: the probability of a yes, the selected option with a distribution over the options, or a score along ordered
levels. Your code stays in control: it asks narrow questions and decides what to do with the answers.
That is why this module has its own client. Jev is not an LLMClient, returns no Completion, and has no conversation and no
streaming. Sending a prompt through a chat client and parsing generated JSON would lose the typed question and answer
contract; so would treating Jev as an OpenAI-compatible provider.
Not yet published.
llm4s-jevexists in the build as of #1265 but ships in the next release.
1
2
// same version as llm4s-core
libraryDependencies += "org.llm4s" %% "llm4s-jev" % llm4sVersion
It brings no dependency beyond llm4s-core, and is a Beta module under modules/providers/jev.
Configure it
Set your TypeSafe API key in TYPESAFE_API_KEY, the variable TypeSafe’s own SDKs read. Nothing else is needed:
1
2
3
4
import org.llm4s.config.JevConfigLoader
// Result[JevClient]: Left(ConfigurationError) when no key is set
val client = JevConfigLoader.default().flatMap(JevClient(_))
JevConfigLoader (in org.llm4s.config, like LLM4S’s other loaders) reads the llm4s.jev block of your
application.conf at the application edge; JevClient itself reads no configuration and takes only the JevConfig it is
given. Every setting has a default:
1
2
3
4
5
6
7
8
9
10
11
12
13
llm4s.jev {
model = "jev-1.13.0" # pin the answers; the default is the alias jev-latest, which moves when a release ships
timeout = 20 seconds # per HTTP attempt; the default is 30 seconds
retry {
maxRetries = 3 # after the first attempt; defaults: 2, 500 ms, 5 s, 0.25, 30 s
backoffInitial = 500 milliseconds
backoffMax = 5 seconds
jitter = 0.25
budget = 30 seconds
}
# apiKey = ... # wins over llm4s.credentials.typesafe.apiKey, which TYPESAFE_API_KEY binds
# baseUrl = ... # TYPESAFE_BASE_URL; default https://api.typesafe.ai
}
| Setting | Environment variable | Default |
|---|---|---|
llm4s.credentials.typesafe.apiKey (or llm4s.jev.apiKey) |
TYPESAFE_API_KEY |
none: required |
llm4s.jev.baseUrl |
TYPESAFE_BASE_URL |
https://api.typesafe.ai |
llm4s.jev.model |
TYPESAFE_DEFAULT_MODEL |
jev-latest |
Or build a config in code:
1
2
3
4
5
6
val config = JevConfig(apiKey = "tsk-...")
.withModel("jev-1.13.0") // pin the answers; the default is the alias jev-latest
.withTimeout(20.seconds)
.withRetry(JevRetryPolicy(maxRetries = 3))
val client = JevClient(config) // validates the config
The key is a bearer token, so the client refuses to send it in the clear: the base URL must be https, except for a loopback
host (localhost, 127.0.0.0/8, ::1), which is how a test points the client at a local server. A URL that carries
credentials, such as http://localhost@evil.example/, is refused. The key never appears in toString or in a log line, and
an error masks it wherever a server echoed it in one of the forms listed under Errors.
Ask questions
A request is a state and a map of questions. All the questions are answered against the one state in a single call, so ask
everything you might need at once and let your code decide which answers matter.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
val decision = for {
client <- JevClient(config)
response <- client.evaluate(
JevRequest(
"Help! My payouts have been failing for 3 days.",
Map(
"department" -> JevQuestion.choice(
"Which team should handle this?",
"billing" -> "Payments, invoicing, refunds",
"technical" -> "Bugs, outages, integrations",
"sales" -> "Pricing, upgrades, new accounts"
),
"urgent" -> JevQuestion.noul("Does this need immediate human attention?"),
"frustration" -> JevQuestion.score("How frustrated is the customer?", "Calm", "Frustrated", "Very angry")
)
)
)
department <- response.choice("department")
urgent <- response.noul("urgent")
frustration <- response.score("frustration")
} yield (department.choice, department.confidence, urgent.probability, frustration.score)
| Question | You give | The answer carries |
|---|---|---|
JevQuestion.Noul |
a yes/no question, optionally what yes and no mean | probability of yes, 0 to 1 |
JevQuestion.Choice |
options, each with an optional description (at most 255) | choice, a probabilities map and a confidence |
JevQuestion.Score |
2 to 10 ordered level descriptions | score (it can land between levels; one a floating-point rounding error, within 1e-9, outside 0 to the highest level is clamped into range), levels with their probabilities, and a confidence |
The ids ("department", …) are yours; the answers come back under the same ids and Jev never sees them. state and
instructions can be a string, or JSON structure when a question refers to data: see TypeSafe’s
API reference. A request that breaks a documented limit is refused with a ValidationError
before anything is sent.
confidence is not the probability of the answer. It says how certain the model is, derived from the whole distribution, and
is what you gate an automatic action on. A Noul has no confidence on the wire. Each response also carries model, the
versioned model that answered (jev-1.13.0 for jev-latest), which you should log, and usage.
A worked example
JevTicketTriageExample
asks the three questions above about a support ticket and routes it with thresholds that are ordinary code: escalate when the
ticket looks urgent, queue it with its department otherwise, and hand it to a person when Jev is not sure which department.
1
TYPESAFE_API_KEY=... sbt "samples/runMain org.llm4s.samples.jev.JevTicketTriageExample"
It is a plain workflow step. Handing a routed ticket to a specialist agent is a Handoff at the point where the route is used,
and needs an LLM provider as well, so the sample stops at the decision.
Errors
evaluate blocks and returns every failure as a Result:
1
2
3
4
5
6
7
8
9
10
def describe(error: LLMError): String = error match {
case _: ValidationError => "the request is wrong: fix it, retrying will not help"
case _: AuthenticationError => "the key was rejected"
case rate: RateLimitError => s"rate limited, retry after ${rate.retryAfter.getOrElse(30.seconds)}"
case svc: ServiceError => s"TypeSafe failed with ${svc.httpStatus}"
case _: TimeoutError => "no answer in time"
case _: NetworkError => "could not reach TypeSafe"
case _: CancelledError => "the calling thread was interrupted"
case other => other.message
}
| What happened | Error |
|---|---|
| The request breaks a documented limit | ValidationError, before anything is sent |
| 401, 403 | AuthenticationError |
| 400, 422 | ValidationError, carrying the server’s explanation |
| 429 | RateLimitError, carrying the delay the server asked for |
408, 5xx including 529 Overloaded, and any other status |
ServiceError |
No connection, or no answer within timeout |
NetworkError, TimeoutError |
| A 200 whose body is not the documented shape, that leaves a question unanswered or answers one that was not asked, or that answers one with another type, an option not offered or a level not described, or that leaves out an option or level that was asked | ProcessingError |
| The calling thread is interrupted | CancelledError, with the interrupt flag kept |
TypeSafe does not document the JSON shape of an error body, so the text in the error is a best effort (a message, an
error.message, …), truncated, and never the whole body. If a server echoes your API key in an error body, the key is
removed before anything is read from it, and an error about a 200 that quotes the key masks it as ***. The forms
recognised are the key as written, JSON-escaped (\/ and \uXXXX included) and URL-encoded as java.net.URLEncoder
writes it (tsk/live+0 as tsk%2Flive%2B0). Any other transformation of the key, or a part or prefix of it, is not
recognised and would not be masked. An error body nested more than 32 levels deep (real ones take a handful of levels,
four for a FastAPI-style detail list) is not read at all: the error carries only the status. A 200 nested more than 64
levels deep is refused as a ProcessingError: the envelope takes four levels, so a structured Score level description
may nest up to 60, and since Jev echoes each description back, a request with one nested deeper is refused with a
ValidationError before it is sent. These limits are well below the 512 levels LLM4S allows elsewhere because the
client walks an error body and keeps a description in the response, and at 512 levels those recursions can overflow a
small thread stack.
Retries, and what is not de-duplicated
Transient failures are retried the way TypeSafe’s own SDKs do (JevRetryPolicy.default): a 408, 429 or 5xx, a connection
failure or a timeout, up to two retries, waiting 0.5 s doubling to at most 5 s with a quarter of each wait randomly taken off,
within a 30 s budget for the whole call. A delay the server asks for (Retry-After, or retry-after-ms, which wins when both
are sent) replaces the computed one, and no retry is started whose wait would reach the budget. A rejected key, an invalid
request and an interrupt are never retried, and the last error is returned as it is.
There is no idempotency key. TypeSafe’s documentation describes no idempotency key, and no other mechanism to de-duplicate a request, so this client does not invent one: each attempt of a retried request is a separate, billable call. If TypeSafe documents a header for it, attach it to the request, and it is sent unchanged on every attempt:
1
2
val request = JevRequest("...", Map("urgent" -> JevQuestion.noul("Is it?")))
.withHeader("X-Correlation-Id", "order-4711")
Headers a client sets itself (Authorization, Content-Type, Accept, …) cannot be replaced, and a header name or value
with a line break is refused. Header names are case-insensitive: a request header replaces a configured one of the same name
in any case, and a map naming one header twice (X-Trace and x-trace) is refused.
Limits and what is not verified
- Jev is text only (a string, or JSON of text values), has no streaming, and its context is 64k tokens per request, 32k for
the
stateplus the longest question. TypeSafe’s rate limits can change without notice. See Models. - The client is tested against a local fake server built from TypeSafe’s published API reference. It has not been run
against the live API. Three details are not in that documentation and are assumptions: the JSON shape of an error body,
whether
Retry-Afteris seconds or a date (both are read), and which ofRetry-Afterandretry-after-mswins when both are sent (this client prefersretry-after-ms). - Each HTTP attempt has a 30 s timeout by default, an LLM4S choice (the API documents none). The shared HTTP client reads a response in full, so the cap on a response (16,777,216 UTF-16 code units, counted once the body is decoded to text) refuses an oversized body before it is parsed, not before it is read.