All blocking LLM calls are shifted to the blocking thread pool via Async[F].interruptible, keeping the compute pool free for fibers. Cancelling the fiber (or a timeout) interrupts the provider thread, which is how llm4s providers are cancelled: they keep the interrupt and return Left(CancelledError). LLMError values are surfaced as LLMException in the F error channel.
The underlying blocking call runs on an interruptible blocking thread and pushes each chunk into a bounded queue (backpressure: the provider thread blocks while the consumer lags). Chunks are delivered as they arrive, not after the call completes. If the call fails mid-stream, chunks already received are emitted first and the stream then fails with LLMException. Stopping consumption early (e.g. take) or cancelling the fiber interrupts the blocking call.