Caches the token fetch returns until refreshMargin before it expires - or half-way through its life, if it lives shorter than twice the margin. Concurrent callers share one in-flight fetch.
A failed fetch is shared for a brief window: the callers waiting behind it get the same failure at once instead of each making their own attempt in turn, which during an outage would hold the last of N callers for N times the exchange's timeout. A rejection (an authentication or configuration error, or any other error not marked recoverable) is shared for FailureTtl (5 seconds); a transient failure (a org.llm4s.error.RecoverableError: network, timeout, rate limit, 5xx) for TransientFailureTtl (1 second) - long enough for the callers queued behind the fetch, short enough that a blip does not outlive itself. After the window the next call tries again; a success, or a rejected cached token, ends it sooner.
A cancellation is never shared. A fetch whose thread is interrupted - it returns CancelledError, throws InterruptedException, or fails while the flag is set - returns CancelledError to that caller alone, with its interrupt flag set, and caches nothing: the next caller through the lock makes its own fetch.
Attributes
- Companion
- object
- Graph
-
- Supertypes