---
title: "Azure DevOps Agent Pools: Fix Idle Agents and Queue Stalls"
description: "Diagnose idle Azure DevOps agents blocking your queue: check parallel jobs, agent demands, stuck leases, and pool scaling before you rebuild a pool."
canonical: "https://adamtheautomator.com/azure-devops-idle-agents-queue-stalls/"
---

# Azure DevOps Agent Pools: Fix Idle Agents and Queue Stalls

> Diagnose idle Azure DevOps agents blocking your queue: check parallel jobs, agent demands, stuck leases, and pool scaling before you rebuild a pool.

Source: https://adamtheautomator.com/azure-devops-idle-agents-queue-stalls/

---

ATA Learning

Tap to hide

[

ATA Learning

](/)

*   [Home](/)
*   [Tutorials](/tutorials/)
*   [Instructors](/author/)
*   [Advertising](/advertising/)
*   [Recommended Resources](/resources/)
*   [About Adam](/about-adam/)

Search for:  

*   [](https://twitter.com/adbertram)
*   [](https://github.com/Adam-the-Automator)
*   [](https://www.linkedin.com/company/adam-the-automator-llc)
*   [](/feed/)

![Azure DevOps Agent Pools: Fix Idle Agents and Queue Stalls](https://adamtheautomator.com/wp-content/uploads/publisher/2e05d9c85b2b819f93afc24ff9cf331a/cc243abf3399435c522e346ae33dfd425022d22b25c2f30ec3f4222e9b893e33.webp)

# Azure DevOps Agent Pools: Fix Idle Agents and Queue Stalls

[![](https://secure.gravatar.com/avatar/d0b9d42e21e5622713f8b693aa5c0f9244d5f7dd200ed29b8398f52dee5de337?s=192&d=mm&r=g)Adam Bertram](https://adamtheautomator.com/author/adam-bertram/)1 October 202613 min. read

Categories: [DevOps](/category/devops/)

Tags:[Azure Pipelines](/tag/azure-pipelines/)[Azure Virtual Machines](/tag/azure-virtual-machines/)[DevOps](/tag/devops/)[Azure Monitor](/tag/azure-monitor/)

Table of Contents

*   [Why Idle Agents and Queued Jobs Exist at the Same Time](#why-idle-agents-and-queued-jobs-exist-at-the-same-time)
*   [The Three Ceilings That Starve a Pool](#the-three-ceilings-that-starve-a-pool)
*   [Read the Queue Before You Touch the Pool](#read-the-queue-before-you-touch-the-pool)
*   [Pull the Numbers the Pool Page Hides](#pull-the-numbers-the-pool-page-hides)
*   [Read the Two Numbers Together](#read-the-two-numbers-together)
*   [Watch the Consumption History](#watch-the-consumption-history)
*   [The Five Failure Modes That Leave a Pool Idle](#the-five-failure-modes-that-leave-a-pool-idle)
*   [Rule These Out In Order](#rule-these-out-in-order)
*   [Demands vs Capabilities: When Nothing Matches](#demands-vs-capabilities-when-nothing-matches)
*   [Capability Changes Need a Restart](#capability-changes-need-a-restart)
*   [Narrow the Pool Instead of Widening It](#narrow-the-pool-instead-of-widening-it)
*   [Diagnosing an Idle Azure DevOps Agent Pool by Pool Type](#diagnosing-an-idle-azure-devops-agent-pool-by-pool-type)
*   [What You Can Read on Each Pool Type](#what-you-can-read-on-each-pool-type)
*   [Turn On Verbose Logging Before You Reproduce](#turn-on-verbose-logging-before-you-reproduce)
*   [Stuck Leases and Orphaned Job Requests](#stuck-leases-and-orphaned-job-requests)
*   [The Evidence Trail](#the-evidence-trail)
*   [When to Escalate](#when-to-escalate)
*   [Scale-In and Idle Sampling: Why You See More Idle Agents Than You Asked For](#scale-in-and-idle-sampling-why-you-see-more-idle-agents-than-you-asked-for)
*   [The Convergence Timeline](#the-convergence-timeline)
*   [The Sampling Window Has a Cost of Its Own](#the-sampling-window-has-a-cost-of-its-own)
*   [The Triage Loop: Diagnostics Tab, Unhealthy Agents, and Saved VMs](#the-triage-loop-diagnostics-tab-unhealthy-agents-and-saved-vms)
*   [Follow the Loop In Order](#follow-the-loop-in-order)
*   [One Saved VM Is the Limit](#one-saved-vm-is-the-limit)
*   [Prevention Playbook: Standby Schedules, Grace Periods, and Maintenance Windows](#prevention-playbook-standby-schedules-grace-periods-and-maintenance-windows)
*   [Size Standby Against Real Demand](#size-standby-against-real-demand)
*   [Give Stateful Agents a Grace Period](#give-stateful-agents-a-grace-period)
*   [Stop Image Updates From Killing Agents Mid-Flight](#stop-image-updates-from-killing-agents-mid-flight)
*   [Check Agent Compatibility Before the Next Auto-Upgrade](#check-agent-compatibility-before-the-next-auto-upgrade)
*   [What Idle Agents Cost You, and What to Do Next](#what-idle-agents-cost-you-and-what-to-do-next)
*   [The Three Moves That End Most Idle-Agent Tickets](#the-three-moves-that-end-most-idle-agent-tickets)
*   [Frequently Asked Questions](#frequently-asked-questions)
*   [Why Do My Agents Show as Idle While Jobs Are Queued?](#why-do-my-agents-show-as-idle-while-jobs-are-queued)
*   [Does an Expired Personal Access Token Take My Agents Offline?](#does-an-expired-personal-access-token-take-my-agents-offline)
*   [Can I Read Agent Logs on a Managed DevOps Pool?](#can-i-read-agent-logs-on-a-managed-devops-pool)
*   [What Does “This Agent Request Is Not Running Yet” Actually Mean?](#what-does-this-agent-request-is-not-running-yet-actually-mean)

Don’t rebuild the pool yet. An Azure DevOps agent pool reporting twelve agents online and idle while your pipeline sits on `This agent request is not running yet. Current position in queue: 4` has already given you the first clue, and the fault usually sits outside those machines. Recreating the agents burns a maintenance window, wipes every capability you installed on them, and lands you back in the same queue position with new hostnames to register.

The order below separates an idle-agent problem from a capacity problem: queue and agent states first, then demands against capabilities, then scaling settings, and only last the host.

## Why Idle Agents and Queued Jobs Exist at the Same Time

Two independent limits decide whether a queued job starts. “Parallel jobs are configured at the Azure DevOps organization level and shared by all pipelines in the organization’s projects,” as the concurrent jobs documentation puts it. One [parallel job](https://learn.microsoft.com/en-us/azure/devops/pipelines/licensing/concurrent-jobs?view=azure-devops) buys one concurrent pipeline job, so an Azure DevOps agent pool can show twenty online agents while the organization holds five.

A pool adds its own ceiling: **Maximum number of virtual machines in the scale set** on scale set pools, and **Maximum agents** on [Managed DevOps Pools](https://learn.microsoft.com/en-us/azure/devops/managed-devops-pools/configure-pool-settings?view=azure-devops). The [Managed DevOps Pools troubleshooting guidance](https://learn.microsoft.com/en-us/azure/devops/managed-devops-pools/troubleshooting?view=azure-devops) does the arithmetic: “if your organization’s self-hosted parallel count is 10, your organization can run only 10 self-hosted pipeline jobs concurrently.”

### The Three Ceilings That Starve a Pool

| Ceiling | Where You Set It | What It Caps | What the Pool Page Shows |
| --- | --- | --- | --- |
| Self-hosted parallel jobs | Organization settings, then [concurrent job limits](https://learn.microsoft.com/en-us/azure/devops/pipelines/licensing/concurrent-jobs?view=azure-devops) | Concurrent self-hosted jobs across every project | Agents online and idle, queue long |
| Maximum agents | Scale set pools and [Managed DevOps Pools configuration](https://learn.microsoft.com/en-us/azure/devops/managed-devops-pools/configure-pool-settings?view=azure-devops) | Agents the pool may run at once | Agent count pinned at the configured number |
| Number of agents to keep on standby | [Scale set pool settings](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/scale-set-agents?view=azure-devops) | Idle agents kept warm for instant assignment | Queue moves, agents arrive late |

An agent is [idle](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/scale-set-agents?view=azure-devops) when it is online and not executing a pipeline job. The definition says nothing about whether it matches the job’s demands.

* * *

_**Reality Check: Idle and online describes an agent’s state. It is not evidence that the pool has capacity to spare.**_

* * *

## Read the Queue Before You Touch the Pool

The wording of the wait reason narrows the search. `All available agents are in use` points at concurrency. `No agent found in pool <pool> which satisfies the specified demands: <demands>` points at a capability mismatch. `Current position in queue: N` points at a request that was never assigned. [Troubleshoot pipeline failure to start](https://learn.microsoft.com/en-us/azure/devops/pipelines/troubleshooting/troubleshoot-start?view=azure-devops) covers all three causes.

### Pull the Numbers the Pool Page Hides

The [Distributed Task REST API](https://learn.microsoft.com/en-us/rest/api/azure/devops/distributedtask/pools?view=azure-devops-rest-7.1) answers the same question with counts. Export your organization and a personal access token (PAT) scoped to [**Agent Pools (read, manage)**](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/personal-access-token-agent-registration?view=azure-devops) so the credential stays out of your shell history.

```bash
export AZDO_ORG="<your-organization>"
export AZDO_PAT="<agent-pools-read-manage-pat>"
```

Azure DevOps agent pool IDs come from that same REST reference. The job requests endpoint returns one record per pending request.

```bash
# Pending job requests for one pool
curl -s -u ":$AZDO_PAT" \
  "https://dev.azure.com/$AZDO_ORG/_apis/distributedtask/pools/<pool-id>/jobrequests?api-version=7.1"
```

The agents endpoint returns each agent’s `status` and `enabled` flag, which is how you spot a ghost record left behind by a virtual machine that was deleted without unregistering.

```bash
# Agent names, online/offline status, and enabled flag
curl -s -u ":$AZDO_PAT" \
  "https://dev.azure.com/$AZDO_ORG/_apis/distributedtask/pools/<pool-id>/agents?api-version=7.1" \
  | jq '.value[] | {name, status, enabled}'
```

### Read the Two Numbers Together

1.  Pending above zero, idle at zero. Capacity problem. Check organization parallel jobs and the pool’s maximum agent setting before touching a host.
    
2.  Pending above zero, idle above zero. Assignment problem. Demands, the parallel job count, or a disabled agent record is blocking the match.
    
3.  Pending at zero, idle above zero. Healthy pool waiting for work. Leave it alone.
    
4.  Offline above zero, pending above zero. Registration or connectivity failed. Start with the `_diag` agent log and the network path.
    

### Watch the Consumption History

The [pool consumption report](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/pool-consumption-report?view=azure-devops) plots online agents, queued jobs, and running jobs at ten-minute granularity. The four patterns above map to one decision, shown below.

![Idle vs queue depth](https://adamtheautomator.com/wp-content/uploads/publisher/d97133409ed3b25f33267c8f6e66b19b2b33069a06c41f69b0e095e21c46435b.png)

## The Five Failure Modes That Leave a Pool Idle

Every idle-agent incident ends in one of five places, ordered here from the checks that take under a minute to the ones that need a shell on a host.

### Rule These Out In Order

1.  Organization concurrency is exhausted. Open **Organization settings**, then **Parallel jobs**, and compare the self-hosted count to the jobs running now. The count is shared by every project.
    
2.  A demand matches nothing. The next section covers this case, because the fix differs from the other four.
    
3.  An agent record is stale or disabled. Delete the stale record before registering a replacement, and give the replacement a unique name. Reusing the name collides with the existing entry and leaves the agent stuck offline.
    
4.  The agent cannot reach the service. Outbound HTTPS on port 443 to the documented [Azure DevOps domains and IP ranges](https://learn.microsoft.com/en-us/azure/devops/organizations/security/allow-list-ip-url?view=azure-devops) is mandatory, and an inspecting proxy breaks registration even when a browser on the same host works. The `_diag` log shows a 60-second timeout against the pool messages endpoint followed by `Agent connect error: The HTTP request timed out after 00:01:00.. Retrying until reconnected.`
    
5.  Credential refresh is looping. The registration token is used only while the agent is configured. “A single PAT can be used for registering multiple agents, the PAT is used only at the time of registering the agent, and not for subsequent communication,” per the agent registration documentation, so an expired registration token does not take a running agent offline.
    

A healthy refresh logs four lines in sequence: `Authentication failed with status code 401.`, `Started authentication`, `acquired new token`, and `Finished authentication`. An agent that repeats the first line alone has a real refresh failure, the pattern in [azure-pipelines-agent issue #5023](https://github.com/microsoft/azure-pipelines-agent/issues/5023). The [log review guidance](https://learn.microsoft.com/en-us/azure/devops/pipelines/troubleshooting/review-logs?view=azure-devops) lists where each diagnostic file lands.

* * *

_**Warning: Deleting an agent record and re-registering under the same name can leave the agent stuck offline. Remove the record first, then register with a new, unique name.**_

* * *

## Demands vs Capabilities: When Nothing Matches

A demand is a rule the job applies to the pool; a capability is a name and value pair an agent advertises. When no online agent satisfies the demand set, the job queues forever while every agent idles. The [pool demands reference](https://learn.microsoft.com/en-us/azure/devops/pipelines/yaml-schema/pool-demands?view=azure-pipelines) documents exactly two demand operations: `Exists` and `Equals`.

```yaml
pool:
  name: linux-agents
  demands:
  - docker          # the agent must advertise a capability named docker
  - Agent.Version -equals 3.227.1
```

### Capability Changes Need a Restart

Installing a tool doesn’t add the capability until the agent process restarts and rescans the machine. The [agent documentation](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/linux-agent?view=azure-devops) is blunt: “After you install new software on an agent, you must restart the agent for the new capability to show up in the pool, so that the build can run.” Pipeline diagnostics write a `capabilities.txt` file beside the other agent logs, which the [Microsoft Learn log review guidance](https://learn.microsoft.com/en-us/azure/devops/pipelines/troubleshooting/review-logs?view=azure-devops) describes as “a clean way to see all capabilities installed on the build machine that ran your build.”

### Narrow the Pool Instead of Widening It

And if only two of your forty agents carry the toolchain a job needs, a demand routes every instance of that job to those two while thirty-eight sit idle. Split the pool, or [push the toolchain into your image build](https://adamtheautomator.com/deploy-enterprise-powershell-modules-using-azure/).

## Diagnosing an Idle Azure DevOps Agent Pool by Pool Type

Diagnostic surfaces differ by pool type. On a [managed pool](https://learn.microsoft.com/en-us/azure/devops/managed-devops-pools/overview?view=azure-devops), the difference is total: no shell on the agent, no `_diag` folder to read.

### What You Can Read on Each Pool Type

| Surface | Self-Hosted | Scale Set | Managed DevOps Pools | Microsoft-Hosted |
| --- | --- | --- | --- | --- |
| Agent `_diag` logs | Yes, on the host | Yes, via the saved unhealthy agent | No host access | No host access |
| Pool **Diagnostics** tab | Limited | Yes | Provisioning events via [Azure Monitor](https://learn.microsoft.com/en-us/azure/devops/managed-devops-pools/diagnostics?view=azure-devops) | Not offered on Microsoft-hosted pools |
| Agent version control | Manual | Managed by the pool | Managed by the pool | Managed by Microsoft with the image |

On a self-hosted agent you own the host, so you read the agent and worker logs straight from the [`_diag` folder](https://adamtheautomator.com/azure-devops-pipeline-down-troubleshooting-guide/) in the agent installation directory. For scale set agents, open **Project settings**, then **Agent pools**, then the pool’s Diagnostics tab, which lists every create, delete, and reimage action with the errors behind them, including subscription quota failures for cores, disks, or IP addresses.

Managed pools deliver the equivalent through the pool’s Azure Monitor diagnostic settings on the pool resource. Those provisioning logs answer why an agent failed to provision, and they carry no job execution data.

Microsoft-hosted agents are the fourth case, with the thinnest diagnostic surface: Microsoft owns the machine, so there is no shell to open and no `_diag` folder to read. Verbose logging through `System.Debug` still applies here. The run’s log download after the job finishes carries the rest of the evidence. The orphaned-request report below describes a Microsoft-hosted incident whose clean-up path runs through Microsoft support.

### Turn On Verbose Logging Before You Reproduce

Verbose logging is the cheapest instrumentation available. It also works on every pool type. Set the pipeline variable `System.Debug` to `true` for all runs, or use **Enable system diagnostics** when you’re queueing one run.

```yaml
variables:
  system.debug: true
```

`System.Debug` also sets `Agent.Diagnostic` to `true`, which collects deeper network and system diagnostics. `Agent.Diagnostic` requires [agent software version 2.200.0 or later](https://learn.microsoft.com/en-us/azure/devops/pipelines/build/variables?view=azure-devops), the version this post targets alongside REST API 7.1. Verbose logs add one resource line per step:

```text
Agent environment resources - Disk: D:\ Available 12342.00 MB out of 14333.00 MB, Memory: Used 1907.00 MB out of 7167.00 MB, CPU: Usage 17.23%
```

The `Agent environment resources` line reports disk, memory, and CPU headroom together. The diagram below shows where the evidence lives on each pool type.

![Diagnostic surfaces](https://adamtheautomator.com/wp-content/uploads/publisher/0ba6bfb7f900168549ae6121c615a1ad7d7865be6226c79a3255a2ed468425fc.png)

## Stuck Leases and Orphaned Job Requests

A lease that never releases holds its parallelism slot while the agent returns to idle. The organization shows idle agents and a queue that will not drain, and pool changes move neither number. A canceled run whose dispatcher never released the lease is the usual origin.

### The Evidence Trail

Support will ask for four items before it can act on an escalation.

1.  `GET .../pools/{poolId}/jobrequests` output showing pending requests while the agents endpoint shows online, idle agents.
    
2.  The run page for each canceled run, with the run ID and cancellation timestamp.
    
3.  The pool consumption report for the same window, showing running jobs below the parallel job count.
    
4.  The organization’s self-hosted parallel job count and the pools that could be holding a slot.
    

### When to Escalate

If orphaned requests survive two cancellation cycles and a pool with idle agents still cannot pull from the queue after thirty minutes, you have a platform-side condition, and restarting agents will not clear a request the dispatcher never released. Confirm the job’s [`cancelTimeoutInMinutes`](https://learn.microsoft.com/en-us/azure/devops/pipelines/yaml-schema/jobs-job?view=azure-pipelines) is long enough for cleanup to finish first: the [pipeline runs documentation](https://learn.microsoft.com/en-us/azure/devops/pipelines/process/runs?view=azure-devops) explains that jobs get “a grace period called the cancel timeout in which to complete any cancellation work.” [A Microsoft Learn Q&A report](https://learn.microsoft.com/en-us/answers/questions/5793414/microsoft-hosted-agents-permanently-stuck-orphaned) documents orphaned requests and every failed attempt to clear them.

## Scale-In and Idle Sampling: Why You See More Idle Agents Than You Asked For

Excess idle agents are often the scaling algorithm working as designed. [The scale set autoscaling documentation](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/scale-set-agents?view=azure-devops) spells out the sampling loop: “Azure Pipelines samples the state of the agents in the pool and virtual machines in the scale set every 5 minutes.”

Scale-out fires when idle agents drop below the standby count, or when waiting jobs exist with zero idle agents. Scale-in fires only after idle agents exceed the standby count for thirty minutes, configurable through **Delay in minutes before deleting excess idle agents**.

### The Convergence Timeline

Plan the on-call rotation around these intervals instead of around a stopwatch.

| Event | Documented Interval | Source |
| --- | --- | --- |
| Pool and scale set sampling | Every 5 minutes | [Scale set agent management](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/scale-set-agents?view=azure-devops) |
| Scale-out step, machines created | Allow 20 minutes for machines to be created for each step | Scale set agent management |
| Scale-in trigger | Idle above standby for more than 30 minutes | Scale set agent management |
| Agent marked unhealthy if it never comes online | 10 minutes | Scale set agent management |
| On-demand provisioning in managed pools | A few moments to 15 minutes | [Managed DevOps Pools scaling](https://learn.microsoft.com/en-us/azure/devops/managed-devops-pools/configure-scaling?view=azure-devops) |

Microsoft’s own framing is blunt: “You may observe more idle agents than you desire at various times, which is expected as Azure Pipelines converges gradually to the constraints that you specify.”

### The Sampling Window Has a Cost of Its Own

The same five-minute sampling window can hold every agent busy and still not scale out. The scale set documentation notes that “Due to the sampling size of 5 minutes, it’s possible that all agents can be running pipelines for a short period of time and no scaling out will occur.” A burst queues jobs even though the pool autoscales.

## The Triage Loop: Diagnostics Tab, Unhealthy Agents, and Saved VMs

Once the cheap checks are done, the remaining work is reconstruction. Turn on the setting that preserves evidence, then follow the loop.

### Follow the Loop In Order

1.  **Enable Save an unhealthy agent for investigation** on the pool. Without it, the failing virtual machine is deleted and the evidence goes with it.
    
2.  **Open the Diagnostics tab** and read the failed actions it logs. Capacity problems surface as quota errors for cores, disks, or IP addresses.
    
3.  **Find the saved agent** under **Agents saved for investigation**, then connect to the VM from the scale set’s **Instances** list in the Azure portal.
    
4.  **Read the logs on the saved VM**, starting with `agent_*.log` for the configuration run and `worker_*.log` for the job that failed.
    
5.  **Run the diagnostics suite** with `./run.sh --diagnostics` on Linux or `.\run.cmd --diagnostics` on Windows. It requires [agent version 2.165.0 or later](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/linux-agent?view=azure-devops).
    
6.  **Check directory permissions** if the suite reports `System.UnauthorizedAccessException: Access to the path '/agent/_diag/Agent_<ts>-utc.log' is denied.` That exception is an access-denied on the agent’s log directory. The documented scale set warning covers the account the instances run under: “It’s OK to create the user and grant it extra permissions, but it should not be the primary administrator, and nothing should depend on the password, as the password will be changed.”
    
7.  **Compare agent versions across the pool** before closing the ticket, and note any version drift for the next upgrade window.
    

### One Saved VM Is the Limit

The setting keeps a single virtual machine, and it saves only a VM where the agent failed to start. When the Diagnostics tab shows `deleting unhealthy machine`, you are looking at a creation failure, and the provisioning log or quota error is your only evidence. The [pools and queues guidance](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/pools-queues?view=azure-devops) lists the permission needed to configure pool maintenance alongside it.

## Prevention Playbook: Standby Schedules, Grace Periods, and Maintenance Windows

The steps above clear an incident. But the settings below keep the next one from reaching your queue.

### Size Standby Against Real Demand

Standby mode is off by default in managed pools, so a queued pipeline waits for provisioning instead of a slot. Set the percentage per image to match actual usage: if 75 percent of jobs run on the Windows image, that image gets a 75 percent standby share. But new pools have no usage history, so automatic prediction is weakest during the first month.

### Give Stateful Agents a Grace Period

A stateless pool hands every job a fresh agent, [the stronger supply-chain position](https://adamtheautomator.com/supply-chain-security-azure-devops/). When you need cache hits across consecutive jobs, run a stateful pool with a grace period so the agent waits for follow-on work. The [`maxAgentLifetime` setting](https://learn.microsoft.com/en-us/azure/devops/managed-devops-pools/configure-scaling?view=azure-devops) defaults to seven days and shuts the agent down at that limit even when jobs keep arriving.

### Stop Image Updates From Killing Agents Mid-Flight

Unattended upgrades restart `walinuxagent.service`. The Azure DevOps backend reads that restart as an agent going offline, and it destroys the VM. The remedy recorded in [that same issue’s unattended-upgrade remedy](https://github.com/microsoft/azure-pipelines-agent/issues/5023) is to stop, disable, and mask `unattended-upgrades` and the `apt-daily` timers on the agent image.

### Check Agent Compatibility Before the Next Auto-Upgrade

Agent software version 5 moves the runtime to .NET 10, and the automatic upgrade does not check whether the host operating system supports it. The version 5 upgrade guidance warns that “If an agent automatically upgrades on an operating system that doesn’t support .NET 10, the upgrade can fail or the agent might not start,” which leaves you a pool full of offline agents after an upgrade nobody scheduled. Run the [compatibility script](https://github.com/microsoft/azure-pipelines-agent/tree/master/tools/FindAgentsNotCompatibleWithAgent) across every self-hosted pool, then upgrade on your own schedule with the [version 5 agent guidance](https://learn.microsoft.com/en-us/azure/devops/pipelines/agents/v5-agent?view=azure-devops).

## What Idle Agents Cost You, and What to Do Next

Cost is the pressure that opened this ticket, so put a number on both sides of the tradeoff: a standby agent is a virtual machine you pay for by the hour, and a queue delay is developer time you pay for in meetings. The table below is illustrative, so substitute your own VM SKU’s hourly rate and size the standby count against your measured peak using the [Managed DevOps Pools scaling settings](https://learn.microsoft.com/en-us/azure/devops/managed-devops-pools/configure-scaling?view=azure-devops).

| Setting | Idle Agents Held | Hourly Cost at $0.10 per VM | Tradeoff |
| --- | --- | --- | --- |
| Standby 0, delete after 30 min | 0 outside bursts | $0.00 | Jobs wait for a VM to boot |
| Standby 2, delete after 30 min | 2 overnight | $0.20 | Two jobs start instantly |
| Standby 6, delete after 120 min | 6 overnight | $0.60 | Bursts absorbed without queueing |

### The Three Moves That End Most Idle-Agent Tickets

1.  Fix the ceiling before you fix the agents. Organization parallel jobs and the pool’s maximum agent count cap that pool no matter how healthy its hosts are.
    
2.  Start with the logs rather than a rebuild. The `_diag` agent log, `capabilities.txt`, and the Diagnostics tab each name a different class of cause.
    
3.  An organization that cannot clear a pending job request through the portal or the API needs Microsoft support, and the four evidence items listed earlier decide how fast that case moves.
    

The queue state usually explains itself before you touch a host.

## Frequently Asked Questions

An idle-agent ticket usually arrives with the same four questions.

### Why Do My Agents Show as Idle While Jobs Are Queued?

Idle means an agent is online and not running a job, so the queued work still needs a demand match. Check organization parallel jobs, the pool’s maximum agent count, and the demand set on the queued job.

### Does an Expired Personal Access Token Take My Agents Offline?

Not by itself. The registration token is used only while the agent is configured, and the agent then runs on a listener token it acquires on its own. An agent that goes offline after a token rotates has a refresh failure, visible in its `_diag` log.

### Can I Read Agent Logs on a Managed DevOps Pool?

No. Managed DevOps Pools run on Microsoft-owned infrastructure, so there’s no `_diag` folder and no shell. Collect evidence from inside the job instead: run your capture or trace tool as a pipeline step and upload the output with [PublishPipelineArtifact@1](https://learn.microsoft.com/en-us/azure/devops/pipelines/tasks/reference/publish-pipeline-artifact-v1?view=azure-pipelines). Provisioning events arrive through the pool’s Azure Monitor diagnostic settings.

### What Does “This Agent Request Is Not Running Yet” Actually Mean?

“This agent request is not running yet” means the queued request was never assigned to an agent, so the problem sits above the individual machine. A queue position that never changes points at a capacity or assignment failure that the job requests endpoint will confirm.

Share this article

[Share on X](https://twitter.com/intent/tweet?url=https%3A%2F%2Fadamtheautomator.com%2Fazure-devops-idle-agents-queue-stalls%2F&text=Azure%20DevOps%20Agent%20Pools%3A%20Fix%20Idle%20Agents%20and%20Queue%20Stalls)[Share on Facebook](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fadamtheautomator.com%2Fazure-devops-idle-agents-queue-stalls%2F)[Share on LinkedIn](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fadamtheautomator.com%2Fazure-devops-idle-agents-queue-stalls%2F)

## Related Posts

![](https://adamtheautomator.com/wp-content/uploads/2026/03/featured_image.webp)

### [Prove Every Artifact: Supply Chain Security in Azure DevOps](/supply-chain-security-azure-devops/)

Learn to implement software supply chain security in Azure DevOps with SBOM generation, artifact signing, dependency scanning, and deployment gate enforcement.

![](https://adamtheautomator.com/wp-content/uploads/2020/01/2576527_0.jpg)

### [Build Real-World Azure DevOps Pipelines for ARM Templates](/azure-devops/)

Create and apply Azure DSC configurations to Azure VMs via ARM templates. Master Azure DevOps pipelines in this step-by-step tutorial.

![](https://adamtheautomator.com/wp-content/uploads/publisher/2e05d9c85b2b8187b503fb0f6ceff47e/e15d54857176da3fe72b1f0692faaa7000367a5fbe6edf5f6a5f1034395fea45.webp)

### [Azure Bicep: Bulletproof Production Deployment Patterns](/azure-bicep-production-patterns/)

Own your Bicep blast radius: pin modules in a private registry, gate pull requests with lint and PSRule, and track production changes with stacks.

## Categories

*   [IT Ops](/category/it-ops/)
*   [Cloud](/category/cloud/)
*   [DevOps](/category/devops/)
*   [Home Ops](/category/home-ops/)
*   [Information Security](/category/infosec/)
*   [Software Development](/category/software-development/)

## Site

*   [Home](/)
*   [Tutorials](/tutorials/)
*   [Instructors](/author/)
*   [Advertising](/advertising/)
*   [Recommended Resources](/resources/)
*   [About Adam](/about-adam/)

Copyright 2026© ATA Learning | [Privacy Policy](/privacy/)
