Programmed I/O: Polling, CPU Utilisation, Worked Examples and GATE Question Patterns

Follow one programmed-input transfer from control write to data read, then calculate how many polls one word costs and what a longer poll period buys.

KnowledgeGate Team

Exam prep & CS education

Updated 7 Oct 20266 min read

Students often remember that programmed I/O uses polling, but mix up who checks readiness, who moves each word and why processor time is wasted. Programmed I/O uses register-by-register transfers, and its utilisation can be calculated from the polling interval, cycles per poll and useful word rate. The simplified programmed-I/O model is an exam abstraction for reasoning about CPU involvement, not a claim that every modern device driver spins in the same way.

Programmed I/O in one precise model

Programmed I/O is a CPU-controlled transfer. The processor executes instructions to read an interface status register, keeps polling until the device is ready, and then reads or writes the data register.

Three actors have distinct jobs:

  • The CPU runs the polling loop and moves each data unit with an instruction or short instruction sequence.

  • The I/O interface exposes status, data and control registers.

  • The device performs the external operation and changes the readiness state.

Polling means repeatedly checking status. Busy waiting means the CPU remains occupied during that wait. With interrupt-driven I/O, an interrupt replaces continuous polling, but the CPU still moves the data. With DMA, a controller moves a block between the device and main memory.

For the topic map (interface registers, memory-mapped versus isolated addressing and DMA cycle stealing in one place), use Input Output Organisation in COA. Programmed I/O needs a different lens: status polling, readiness-detection delay, and CPU cycles per useful word.

Do not mix this classification with register addressing. Memory-mapped versus isolated I/O describes how the CPU addresses interface registers. Programmed, interrupt-driven and DMA I/O describe how transfer and control are organised. The register-address side connects naturally to Addressing Modes and Instruction Formats.

Programmed I/O polling, register by register

Assume this input interface: STATUS = 0xB400, DATA = 0xB401 and CONTROL = 0xB402. Bit 0 of STATUS is READY. Writing 0x01 to CONTROL starts an input operation; after READY becomes 1, reading DATA returns 0x4B. These addresses and register conventions define this example, not universal hardware.

Code
write 0x01 to CONTROL
repeat:
    status = read STATUS
until (status & 0x01) != 0
value = read DATA
continue the program

The CPU polls at 0, 40, 80, 120, 160, 200 and 240 ns. The device changes READY from 0 to 1 at 214 ns. Therefore the first six status reads see 0, the seventh sees 1, and the CPU reads 0x4B at 240 ns.

The observed readiness-detection delay is 240 - 214 = 26 ns. With a 40 ns polling interval, the delay lies from 0 up to, but not including, 40 ns in this simplified timing model. That delay is separate from the time needed to produce or transfer data.

Worked example: polling cost and useful throughput

Suppose an 800 MHz CPU services a device delivering 250,000 bytes/s. Each transfer moves one 2-byte word. One polling iteration, including the successful ready check and data-register read, costs 25 cycles. The CPU spins continuously and does no other work between words. These rates and cycle costs define this example, not universal hardware.

Calculate from the device rate first:

  1. Word rate = 250,000 bytes/s / 2 bytes/word = 125,000 words/s.

  2. Arrival interval = 1 / 125,000 words/s = 8 microseconds/word.

  3. At 800 MHz, cycles per word = 8 microseconds/word x 800,000,000 cycles/s = 6,400 cycles/word.

  4. Polls per word = 6,400 cycles/word / 25 cycles/iteration = 256 iterations/word.

  5. Poll rate = 125,000 words/s x 256 iterations/word = 32,000,000 iterations/s.

  6. CPU cost = 32,000,000 iterations/s x 25 cycles/iteration = 800,000,000 cycles/s, which is 100% of an 800 MHz CPU.

A 2,500-byte block arrives in 2,500 bytes / 250,000 bytes/s = 0.01 s = 10 ms and contains 2,500 bytes / 2 bytes/word = 1,250 words. The loop keeps pace with the device while leaving effectively no CPU availability for other work.

The one knob programmed I/O actually has: the poll period

For the side-by-side CPU-occupancy comparison of programmed I/O, interrupts and DMA over one common device window, use Computer Interfaces in COA: Registers, Handshaking and I/O Mode Numericals.

This post never leaves the polling loop.

Keep the 800 MHz CPU, 250,000 bytes/s device and 2-byte words. Instead of spinning, the loop polls once every 4 microseconds. These values and the 25-cycle polling cost define this example, not universal hardware.

  1. Polls per word = 8 microseconds/word / 4 microseconds/poll = 2 polls/word.

  2. CPU cost per word = 2 polls/word x 25 cycles/poll = 50 cycles/word.

  3. Per second = 125,000 words/s x 50 cycles/word = 6,250,000 cycles/s, which is 6,250,000 cycles/s / 800,000,000 cycles/s x 100 = 0.78125% of the CPU.

The cost falls by a factor of 128, and the price is that worst-case readiness-detection delay rises from just under 40 ns to just under 4 microseconds.

That is the whole programmed-I/O design space in one line: poll period buys CPU time and sells response latency, and nothing else about the mechanism changes.

Where programmed I/O makes sense, and where it fails

Programmed I/O is reasonable when a device becomes ready almost immediately, transfers are rare or tiny, firmware is in an early boot phase, or a simple microcontroller can afford to wait. Its advantages are simple control flow and low setup complexity, not better CPU utilisation.

As device latency or transfer volume grows, repeated status reads consume execution slots, use power and stop useful CPU work. In the trace, the 26 ns detection delay gives quick response, but continuous checking pays for that response.

A useful conceptual rule is: choose programmed I/O for simple, short and infrequent waits; interrupts for sporadic events when the CPU should work between them; and DMA for block transfers where per-word CPU handling dominates. Actual system design still depends on its hardware and latency requirements.

Programmed I/O traps and the GATE question pattern

Fix these four traps before attempting questions:

  • Simpler programmed I/O is not automatically faster.

  • High device throughput does not imply good CPU utilisation.

  • An interrupt announces readiness, but does not by itself transfer the word to memory.

  • Memory-mapped I/O is not a synonym for programmed I/O.

Also separate polling latency from transfer time. The trace detected readiness after 26 ns, while the throughput example receives one word every 8 microseconds. They measure different events.

The official GATE 2026 CS syllabus places I/O interface (interrupt and DMA mode) under Computer Organization and Architecture. It gives no topic-wise marks allocation or occurrence count. Programmed I/O remains the useful baseline against which those modes are compared.

The archived GATE 2024 CS Set 1 paper, Question 15, asks for the false claim among statements about cycle-stealing DMA, burst-mode throughput, CPU utilisation under programmed versus interrupt-driven I/O, and vectored interrupts. The claim that programmed I/O gives better CPU utilisation is false: polling occupies the CPU while it waits, whereas interrupts let it execute other work between ready events.

Use three quick checks on this pattern. In interrupt-driven I/O, identify the CPU as the per-word mover. For utilisation, multiply events per second by cycles per event, then divide by CPU cycles per second. For a large block, prefer DMA when avoiding repeated CPU data moves matters.

The short version and the next useful step

  • The CPU polls the status register.

  • The CPU moves each word.

  • Busy waiting can detect readiness quickly but leave poor CPU availability.

  • Interrupts remove continuous polling, while DMA removes most per-word CPU work.

The obvious next question - at what event rate interrupts actually beat polling, and what the CPU saves and restores when they do - is answered in Interrupt-Driven I/O Explained: CPU Handshake, Worked Timing Example and GATE Traps.

For structured GATE CS concept revision, use GATE Guidance by Sanchit Sir, then test mode comparisons and utilisation arithmetic through the GATE Test Series.