Human Review

Pause an agent’s turn for a person: to approve a tool call, answer a question, review an answer, or look at what a step did before the next one runs.

Table of contents

  1. Four ways a turn waits
  2. Static breakpoints
    1. On any graph
  3. Questions from middleware
    1. Reviewing what guardrails refuse
  4. From Java and Kotlin

Four ways a turn waits

A turn that waits for a person ends AgentStatus.Suspended. Everything it waits for is an interrupt with its own InterruptId, in one of four lists:

List Raised by Answer with
approvals a tool’s ToolOutcome.NeedsApproval, or a tool wrapper such as ApprovalMiddleware result.approve(id), reject(id, reason), edit(id, arguments)
questions a tool extending AgentTool.Asking[A, Q, Ans] result.reply(id, answer), the tool’s Ans
middlewareQuestions a middleware extending AgentMiddleware.Asking[Q, Ans] result.reply(id, answer), the middleware’s Ans
breakpoints a static breakpoint on one of the agent’s nodes result.proceed(id)

agent.resume(threadId, answers) takes any non-empty subset of them. The ones you do not answer stay pending, so the result can be Suspended again with what is left: this is incremental approval. An unknown id is refused and the thread is unchanged, as is an approval decision, a middleware answer or a breakpoint answer that does not decode as the asker’s type. A tool question’s answer is different: the tool decodes it, so one that does not decode as the tool’s Ans is accepted, and becomes that call’s error result (Invalid answer for '<tool>': ...), which the model sees. While a thread waits, run and recover on it are refused with PendingInterrupts.

1
2
3
4
5
6
7
agent.run(threadId, "Deploy the release").flatMap { result =>
  result.status match
    case AgentStatus.Suspended(approvals, questions, asked, held) =>
      val answers = approvals.map((id, _) => result.approve(id)) ++ held.map((id, _) => result.proceed(id))
      agent.resume(threadId, answers.toMap) // questions and middleware questions stay pending
    case _ => Right(result)
}

The waiting thread is persisted before run returns, under every durability mode, so another process with the same agent and checkpointer can resume it.

Static breakpoints

A breakpoint holds a task of one of the agent’s nodes, before it runs or after it ran:

1
2
3
4
5
6
val agent = Agent
  .builder("assistant", client)
  .withTools(tools)
  .withInterruptBefore(AgentNode.Tool)   // each tool call, before it runs
  .withInterruptAfter(AgentNode.Model)   // each model call, after its message is stored
  .build()
AgentNode Node Before After
Model <id>/model the model is not called yet the model’s message is stored; its tool calls, or the answer’s guarding, wait
Tool <id>/call-tool the call has not run its result is recorded; the next model call waits for it
Finish <id>/finish the answer is stored; its afterAgent hooks (guardrails) have not run the answer is guarded; the turn ends once answered

Each held task is an interrupt of its own. The tool calls of one message are held one per call, so a reviewer can let them through one at a time. AgentStatus.Suspended.breakpoints lists them as BreakpointRequest(node, phase, call); call is the tool call, for a call held before it runs. result.proceed(id) continues one: a task held before it runs is run, and the same breakpoint does not hold it again; a task held after it ran is not run again - what follows it is released. A breakpoint takes no other answer.

A step continued after a review runs at another node than the step itself: a model call after a wrapModelCall question at <id>/asked/<middleware>/wrapModelCall; a tool call after an approval or a question at <id>/approval, <id>/ask/<tool> or <id>/asked/<middleware>/wrapToolCall; an answer guarded again after an afterAgent question at <id>/asked/<middleware>/afterAgent (the root’s, for a root middleware asking about a handoff target’s answer). A breakpoint after Model, Tool or Finish holds those nodes too, so nothing the step does escapes it: a model message the gate let through still waits before its tool calls run, and a result recorded after an approval still waits before the model sees it. BreakpointRequest.node names the node the step ran at. A breakpoint before holds only the step’s own node: a continued step was just reviewed.

Each task held after it ran is released in a superstep of its own, which replays its routes without running the node again. That superstep counts against the run’s RunBudgets.maxSupersteps, so a run with many held steps may need a larger budget.

Breakpoints are not part of the agent’s graph version, like streaming: a thread an agent holds can be continued by the same agent built without them. A run can add its own with a RunConfig, naming the node ids: agent.run(threadId, query, RunConfig(interruptBefore = Set(NodeId("assistant/model")))). A node the agent does not have is refused with a ValidationError before the thread is touched.

On any graph

Breakpoints belong to the graph runtime, not only to agents. A RunConfig’s interruptBefore and interruptAfter name nodes of any compiled graph; each start, resume and recover brings its own. A held task is parked like a task that returned NodeResult.Suspend, keyed by its task id, and the run pauses at the end of the superstep - its siblings run and commit. RunResult.Suspended reports it as a PendingInterrupt with breakpoint = Some(BreakpointPhase.Before) (its question is the task’s input) or Some(BreakpointPhase.After) (the node’s update is committed; its routes, static edges and join arrivals wait). Answer it with Breakpoint.proceed:

1
2
3
4
5
6
val config = RunConfig(interruptAfter = Set(NodeId("plan")))
for
  held   <- runtime.start(threadId, graph, input, config).flatMap(_.await())
  resumed <- runtime.resume(threadId, graph, Map(InterruptId("0.0") -> Breakpoint.proceed), config)
  done   <- resumed.await()
yield done

Cancelling a run is unaffected: a cancelled run leaves its thread for recover, which holds a task the breakpoint would hold, and runs a continuation that already passed its breakpoint without holding it again.

Questions from middleware

A middleware that needs a person - to review an answer, to edit it, or to supply missing information - extends AgentMiddleware.Asking[Q, Ans], as an asking tool extends AgentTool.Asking. A hook asks by returning ask(question) (from beforeAgent, afterAgent or wrapModelCall) or askAbout(question) (from wrapToolCall), and reads its answer with answered(context) when it runs again:

1
2
3
4
5
6
7
8
final case class Need(what: String) derives ReadWriter
final case class Info(value: String) derives ReadWriter

final class Region extends AgentMiddleware.Asking[Need, Info]:
  val id = MiddlewareId("region")
  override def beforeAgent(input: String, context: RunContext): Result[String] =
    if input.contains("region") then Right(input)
    else answered(context).fold(ask(Need("region")))(r => Right(s"$input (region ${r.answer.value})"))

The turn suspends with a MiddlewareQuestionRequest - the agent and middleware that asked, the hook, and the question as JSON (read it with ToolLoop.middlewareQuestion[Need](request)). Once answered with result.reply(id, Info("eu")), the asking task runs again from its start, and the whole stack of that hook runs again from the outermost middleware; nothing about a stack’s position is stored. What runs again depends on the hook:

Hook While it waits Once answered
beforeAgent nothing of the turn is stored the turn’s input runs again
wrapModelCall nothing of the model call is stored, but a call the wrapper made before asking is counted in usage the model step runs again: a wrapper that asked after calling next calls the model again, unless it returns a completion of its own
wrapToolCall the call has no result the call runs again: its arguments are checked again, with the approval it had
afterAgent the answer is stored, the turn has no outcome yet the answer is guarded again

Answers given earlier in the same stack run are kept while it runs again, so two asking middleware in one stack each ask once, and a tool call’s answer is kept through a later approval or tool question. Because the stack runs again, a middleware outside the asking one runs twice and must not depend on running once. A tool wrapper’s questions are one per call, so the calls of one message can be answered one at a time. A middleware that asks while the tool continues after its own question gets an error result instead, since running the call again would lose the tool’s answer.

Reviewing what guardrails refuse

GuardrailReviewMiddleware(input, output) runs guardrails as GuardrailMiddleware does, but a refusal asks a reviewer instead of blocking. The question is a GuardrailReview(phase, guardrail, reason, text); the answer a GuardrailVerdict: Allow lets the text through, Edit(text) replaces it, and Block blocks the turn as GuardrailMiddleware would have (AgentStatus.Blocked). A verdict holds for the text and phase it was given for, whatever the refusal’s wording: once given, the guardrails do not run on that text again, so an LLM judge whose reason varies from call to call is asked about once and is not called again.

From Java and Kotlin

JAgentStatus.pending() lists every interrupt as a PendingInterrupt, approvals first, then tool questions, middleware questions and breakpoints. kind() is the Java enum InterruptKind: APPROVAL, QUESTION, MIDDLEWARE_QUESTION or BREAKPOINT. A field that only some kinds have is an Optional: toolName() and argumentsJson() when a tool call waits; reason() for an approval; questionJson() for a question; middleware() for a middleware question; node() and phase() (the Java enum BreakpointPhase, BEFORE or AFTER) for a breakpoint. Answer.proceed(id) continues a breakpoint; Answer.reply(id, json) answers either kind of question. An agent with breakpoints or asking middleware is built with Agent.builder and wrapped with Llm4s.wrapAgent. See Suspended turns from Java and Kotlin.