Skip to content

Choosing Which Ready Agent Goes First

In congested agent systems, readiness is eligibility rather than an automatic release command: keeping some ready turns upstream can preserve workflow-level control over ordering, while current evidence shows that priority policy accounts for most of the measured tail-latency benefit.

A software agent finishes a tool call. The result is back. Its next model turn has everything it needs to run.

Feng and colleagues describe a common runtime default: once the turn is ready, the runtime sends it to the inference engine.

Their September 2026 preprint, Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows, inserts one more state into that sequence. A turn can be ready without yet being submitted. The distinction sounds bureaucratic until many agents are sharing the same GPUs.

Under light load, it barely matters. Under heavy load, it can matter a great deal.

The researchers replayed real software-engineering agent traces while holding the workflows, arrival times, model work, and recorded tool delays fixed. The variable was the release policy. In the largest reported improvement, on a SWE-bench trace served by Llama-3.3-70B, the 95th-percentile workflow completion time fell from 1,986 seconds with eager release to 568 seconds with their scheduler: a 71.4 percent reduction, or a 3.5× speedup.

But the most useful lesson is not that systems should make ready work wait. The paper’s own ablations make that claim too simple.

The more durable lesson is that ready and committed are different states. A ready turn can still be compared with other work. Once it is handed to the inference engine, that choice may be gone.

The queue before the queue

An agent workflow is not one model request. It is a chain: model call, tool call, model call, perhaps another tool, and so on. Each finished step can expose a new turn that could run next.

That creates two scheduling layers. The inference engine sees requests that have already been submitted. A workflow runtime can have richer context: which larger job each turn belongs to, how long that job has been waiting, and what other turns are now ready. In Feng and colleagues’ scheduler, it also estimates how much work a candidate turn is likely to require.

Feng and colleagues call submitted but unfinished requests committed turns. Their committed work measure is the sum of stored work estimates for those turns. In the architecture they study, a ready turn that remains outside the engine stays under the workflow scheduler’s control. A committed turn does not. It may be queued, batched, or running inside the engine, but the workflow layer can no longer withdraw it and reconsider whether another workflow should have gone first.

This is the paper’s important reframing. Readiness answers a dependency question: could this turn run? Release answers a resource-allocation question: should this turn enter the shared bottleneck now?

Eager release collapses those questions into one event.

When the engine has spare capacity, collapsing them is nearly free. Across the six lowest-load conditions in the study, the eager and controlled policies had almost identical P95 workflow completion times; their ratios ranged from 0.99× to 1.00×. A ready turn could be sent immediately because there was little competition for what came next.

Congestion changes the value of the intermediate state. If ten turns are ready and the engine is already saturated, submitting all ten does not create ten new opportunities to make a good decision. It can instead move ten decisions out of the layer that knows the workflows best.

A turn waiting upstream is not necessarily idle work. It can be an unspent scheduling option.

What produced the 3.5× result

That phrase—unspent scheduling option—can also overstate the case if it is allowed to carry the headline number by itself.

The proposed scheduler changes two things at once. First, it chooses which ready turn to release using a priority that considers workflow delay risk and estimated turn cost. Second, it controls how much released-but-unfinished work is allowed to accumulate, tightening that budget as congestion grows.

The paper’s ablation separates those mechanisms at one high-load setting with Llama-3.3-70B. The split is revealing.

With FIFO ordering but an adaptive committed-work budget, the scheduler was 1.42× faster than eager release on the SWE-bench trace and 1.10× faster on SWE-Gym. That is evidence that limiting commitment can help.

But ordering contributed much more. Holding the adaptive budget fixed and replacing FIFO with the paper’s tail-aware ordering reduced P95 by 59.5 percent on SWE-bench and 63.5 percent on SWE-Gym. Then, with that better ordering already in place, making the work budget adaptive rather than fixed reduced P95 by another 10.7 and 11.6 percent.

So the 3.5× result is not a demonstration that queue limits alone rescue agent systems. Most of the measured improvement in that ablation came from choosing better among eligible turns. The release boundary matters because it leaves someplace for that choice to happen.

That is a subtler design principle than “wait more.” It is closer to: do not surrender a scheduling decision before the layer with the useful context has made it.

Agent serving is moving up a level

This idea did not appear in an empty field. The novelty boundary is important because the broad lesson—schedule compound jobs as compound jobs—has been arriving from several directions.

At NSDI 2026, Agentix treated agent programs as first-class schedulable objects rather than unrelated model calls. It intercepts calls and gives its scheduler program-level context; across the workloads reported by its authors, it delivered four to fifteen times more program throughput at the same latency than serving baselines including vLLM.

Another 2026 system, CONCUR, found a different reason not to admit every runnable agent at once. Long-lived agent workloads could thrash the GPU’s key-value cache under high concurrency. CONCUR uses feedback from cache pressure to regulate how many agents are active, reporting throughput gains of up to 4.09× on Qwen3-32B and 1.9× on DeepSeek-V3.

Neither result reproduces Feng and colleagues’ experiment. Agentix is evidence that program-level context can be useful to a serving scheduler. CONCUR is evidence that over-admission can hurt for a different resource reason. Together they make the new paper’s narrower contribution easier to place: it makes the ready-to-release handoff itself a first-class control point.

A third system shows why the objective matters as much as the mechanism. SAGA reports a 1.64× geometric-mean improvement in agent task completion time from workflow-aware scheduling on a 64-GPU cluster. It also reports roughly 30 percent lower peak throughput than throughput-optimized batching.

That is not a defect hidden in a footnote. It is the tradeoff.

A scheduler can make compound jobs finish sooner while moving fewer total tokens through the hardware. If the product objective is interactive completion time, that may be exactly right. If the objective is maximum batch throughput, it may not be. “Keep the GPU busy” and “finish the most valuable workflows quickly” are different optimization problems.

The result has a narrow address

The new study is unusually clean in one way: its paired replay keeps the work fixed while changing release decisions. It is also limited in ways that matter for interpreting the number.

Each trace contains 100 workflows, and each plotted point comes from one paired run using a single fixed Poisson-arrival seed. The authors note that an empirical P95 over 100 workflows is determined by a small subset of those jobs. The direction repeats across two trace sets and three model configurations, but there are no multi-seed confidence intervals for the headline P95 differences.

The experiment is also a replay, not a live agent deployment. Recorded tool delays and model work make the systems comparison controlled, but they remove a possible cost of waiting: the world can change while a turn sits ready. A browser session can expire. A file can change. A remote service can return different data. The paper does not establish what delayed release does to those semantic outcomes.

And the boundary is partly architectural. The claimed loss of control occurs because, in the tested design, the workflow runtime cannot reprioritize or cancel a turn after it has been submitted to the engine. A serving engine that accepted live workflow priorities and allowed submitted work to be reordered could preserve some of the same optionality downstream. In such a system, holding turns in a separate upstream queue might matter less.

Most importantly, P95 workflow latency is not the same thing as useful software output. The study does not measure whether faster workflows produce more accepted patches, less reviewer rework, better fairness between tenants, or more durable code. Those are separate questions.

This makes the finding smaller than the slogan “starting work can be bad.” It also makes it more actionable.

A state worth naming

Agent runtimes already know when a dependency has been satisfied. The experiment suggests that they may also need to preserve the state that comes immediately after: eligible, but not yet committed.

Once that state exists, it can be measured. How long did a turn remain ready before release? How much work was already committed when it became ready? Which workflow information changed its priority while it waited? Did better tail latency cost throughput or fairness? Did waiting change the eventual artifact?

Those questions are hard to ask when “ready” is implemented as a callback whose next line submits a request.

The practical consequence is not a mandate to slow agents down. At low load, the study found almost no reason to wait. The consequence is to stop treating dependency completion as a release command when resources are contested.

A turn can be fully ready and still benefit from one more decision.

Sources