Level 8 · Chapter 2

Scheduling

How Windows decides which thread actually runs next — and why CPU usage numbers alone can't tell you whether a system is responsive.

Kernel mechanisms covered what's legal for kernel code to do at any given moment. This chapter is about a related but separate question: given several runnable threads, which one actually gets the CPU right now, and for how long.

Priority and state, not user intent

The dispatcher — the scheduling component tracking every thread's state — chooses which runnable thread runs next based on priority and state, not based on any notion of what the user considers important. A thread is, at any moment, in one of several states: running (actually executing on a CPU right now), ready (runnable, waiting for a CPU to become available), or waiting (blocked on something — I/O, a synchronization object, a timer — and not eligible to run at all until that wait is satisfied). Only threads in the ready or running state are candidates for the scheduler's attention; a waiting thread, no matter how "important" the application it belongs to might seem, simply isn't part of the decision until it becomes ready again.

Quantum: a bounded slice, not a guarantee

A running thread doesn't run indefinitely — it's granted a quantum, a bounded time slice, after which it can be preempted to let another ready thread of equal or higher priority run. This is what keeps a single CPU-bound thread from starving every other runnable thread indefinitely, and it's a large part of why a system with many competing threads still generally stays responsive rather than having whichever thread happened to start first simply monopolize the processor.

Priority boosts: why the model isn't purely static

Windows doesn't treat priority as a perfectly static number applied uniformly forever — certain situations trigger temporary priority boosts: a thread that was waiting on I/O and just became ready again commonly gets a short-lived boost, on the reasoning that a thread which was recently blocked (and is therefore likely still relevant to whatever the user is actively doing) should get a chance to run promptly rather than waiting behind a long queue of already-running CPU-bound work. These boosts decay over time, gradually returning the thread to its normal, configured priority — a deliberate mechanism for improving perceived responsiveness without permanently distorting the overall priority scheme.

Why "high CPU" and "responsive" aren't opposites, or synonyms

This is the single most practically useful correction this chapter can offer: CPU usage percentage alone does not tell you whether a system, or a specific application, is responsive. A UI application can feel completely frozen while showing near-zero CPU usage — because its UI thread is blocked, waiting on I/O or on another thread, not doing any computation the scheduler would even have an opportunity to grant CPU time to. Conversely, a system pegged at 100% CPU across every core can still feel perfectly responsive to interactive use, if the scheduler is correctly prioritizing foreground, interactive work over background CPU-bound tasks. Diagnosing "this feels slow" correctly requires looking at thread state (is the relevant thread actually runnable, or is it waiting on something), not just aggregate CPU percentage.

A worked example: a UI app that feels frozen with low CPU usage

An application whose main thread is synchronously waiting on a slow network call, or on a lock held by another thread that's itself stuck, will appear completely unresponsive to the user — no window repaints, no input handled — while Task Manager shows that application using essentially no CPU at all. This isn't a scheduling failure; the scheduler is working exactly as designed, correctly not granting CPU time to a thread that isn't currently runnable. The actual problem lives in why that thread is stuck waiting, a question scheduling itself can't answer.

A common mistake

Reading "high CPU usage" as inherently a problem, or "low CPU usage" as inherently evidence of a healthy, responsive system, both miss what the scheduler is actually managing. High CPU usage from correctly-prioritized background work is often perfectly fine; low CPU usage from a critical thread stuck waiting on a blocked resource is often the real, invisible-to-that-one-metric problem.

Where this connects

  • Processes & threads covers the structures (ETHREAD and its scheduling-relevant fields) the dispatcher is actually operating on.
  • Kernel mechanisms covers the IRQL rules that constrain what a thread — including the scheduler's own dispatch code — is allowed to do at any given moment.