Pipelining Basics MCQs: 10 Solved Questions with Explanations

Solve ten pipelining questions step by step, including timing calculations, a unit trap, weighted speedup, a non-uniform stage schedule and data interlocks.

KnowledgeGate Team

Exam prep & CS education

Updated 19 Sep 20267 min read

Single-instruction latency, steady-state throughput, fill-and-drain time, unequal stages and dependency stalls require different reasoning. Start every numerical by identifying the slowest stage, latch delay, instruction count and units. In GATE preparation, solve each option set before reading the bold answer, and carry every unit through the working.

1. The 60-second pipelining toolkit

For k stages and n independent items with no stalls, total cycles are k + n - 1. The clock period is max(stage delay) + pipeline-register delay, and total time is cycles multiplied by the clock period. These formulas assume a constant clock and no bubbles.

A 5-stage pipeline needs 5 cycles for its first instruction. It then ideally completes one instruction per cycle, so 10 instructions take 5 + 10 - 1 = 14 cycles, not 10 or 50. Pipelining normally improves throughput, not the latency of one instruction.

Three losses matter. Unequal stage delays make every stage follow the slowest stage, dependencies cause data hazards, and shared hardware causes structural hazards. Read Pipelining in Computer Architecture Explained for the complete concept lesson.

2. Questions 1-2: clock period, pipeline fill and total time

Question 1: five stages and 100 independent instructions

A five-stage pipeline has stage delays of 150,120,150,160 and 140 nanoseconds. The registers that are used between the pipeline stages have a delay of 5 nanoseconds each. The total time to execute 100 independent instructions on this pipeline, assuming there are no pipeline stalls, is _______ nanoseconds.

Options: none; this is a numerical-answer question.

Answer: 17160 nanoseconds. The slowest stage takes 160 ns, so the clock period is 160 + 5 = 165 ns. The pipeline needs 5 + 100 - 1 = 104 cycles. Therefore, total time is 104 x 165 = 17,160 ns. Do not sum all five stage delays for every instruction because the instructions overlap in execution.

Question 2: four stages and 1,000 data items

A 4-stage pipeline has the stage delays as 150, 120, 160 and 140 nanoseconds respectively. Registers that are used between the stages have a delay of 5 nanoseconds each. Assuming constant clocking rate, the total time taken to process 1000 data items on this pipeline will be

  • A. 120.4 microseconds

  • B. 160.5 microseconds

  • C. 165.5 microseconds

  • D. 590.0 microseconds

Answer: C. 165.5 microseconds. Here, Tclk = 160 + 5 = 165 ns, and the cycle count is 4 + 1000 - 1 = 1003. Thus, 1003 x 165 = 165,495 ns = 165.495 microseconds, which rounds to 165.5 microseconds.

3. Questions 3-4: catch a unit mismatch and combine unequal improvements

Question 3: the unit-conversion trap

A 4-stage pipeline has the stage delay as 150,120,160 and 140 ns respectively. Registers that are used between the stages have delay of 5 ns. Assuming constant clocking rate, the total time required to process 1000 data items on this pipeline is

  • A. 160.5 ms

  • B. 165.5 ms

  • C. 120.5 ms

  • D. 590.5 ms

Answer: none of the listed options. The calculation is 1003 x 165 ns = 165,495 ns, which converts to 0.165495 ms. Option B would match only if the stated delays were microseconds, since 165,495 microseconds = 165.495 ms. Carry the unit through every multiplication before choosing an option.

Question 4: weighted overall speedup

In an enhancement of a design of a CPU, the speed of a floating point unit has been increased by 20% and the speed of a fixed point unit has been increased by 10%. What is the overall speedup achieved if the ratio of the number of floating point operations to the number of fixed point operations is 2:3 and the floating point operation used to take twice the time taken by the fixed point operation in the original design?

  • A. 1.155

  • B. 1.185

  • C. 1.255

  • D. 1.285

Answer: A. 1.155. Let the original fixed-point time be f, so one floating-point operation takes 2f. The original time is 2(2f) + 3(f) = 7f. The new time is 2(2f/1.2) + 3(f/1.1) = 200f/33. Overall speedup is 7f / (200f/33) = 231/200 = 1.155.

4. Question 5: schedule a non-uniform four-stage pipeline

Question 5: eight jobs across four stages

Consider a 4 stage pipeline processor. The number of cycles needed by the four instructions I1, I2, I3, I4 in stages S1, S2, S3, S4 is shown below:

Instruction

S1

S2

S3

S4

I1

2

1

1

1

I2

1

3

2

2

I3

2

1

1

3

I4

1

2

2

2

What is the number of cycles needed to execute the following loop? For (i=1 to 2) {I1; I2; I3; I4;}

  • A. 16

  • B. 23

  • C. 28

  • D. 30

Answer: B. 23. Treat I1, I2, I3, I4, I1, I2, I3, I4 as eight jobs moving through four stages. For job j and stage s, use C[j,s] = max(C[j-1,s], C[j,s-1]) + p[j,s]. The stage-4 completion times are 5, 10, 13, 15, 16, 18, 21, 23, so the final instruction completes at cycle 23. Option 16 ignores stage conflicts, while 28 and 30 miss legal overlap.

Four-stage pipeline schedule for the eight-job I1-I4 loop, ending at a 23-cycle makespan.

5. Questions 6-7: why performance falls and how ideal completion works

Question 6: three causes of lost throughput

The performance of a pipelined processor suffers if

  • A. the pipeline stages have different delays

  • B. consecutive instructions are dependent on each other

  • C. the pipeline stages share hardware resources

  • D. All of the above

Answer: D. All of the above. Unequal delays set the clock by the slowest stage and leave faster stages idle. A dependency can insert a data-hazard stall. Shared hardware can insert a structural-hazard stall.

Question 7: ten instructions in a classic five-stage pipeline

A CPU has a 5-stage pipeline with the following stages Fetch (F), Decode (D), Execute (E), Memory (M) and Write-back (W). Each stage takes one clock cycle to complete. Assume there are no pipeline stalls and the pipeline is initially empty. How many clock cycles are required to complete the execution of 10 instructions?

  • A. 10

  • B. 14

  • C. 15

  • D. 19

Answer: B. 14. The first instruction completes after 5 cycles, and the other 9 ideally complete one per cycle: 5 + 9 = 14. Equivalently, k + n - 1 = 5 + 10 - 1 = 14. Latency is 5 cycles/instruction, while steady-state throughput is 1 instruction/cycle.

6. Questions 8-10: operation order, interlocks and the core definition

Question 8: arithmetic-pipeline order

The right sequence of suboperations that are performed in arithmetic pipeline is –

A. Align the mantissas

B. Add or subtract the mantissas

C. Normalize the result

D. Compare the exponents

Choose the correct answer from the options given below:

  • A. D, A, B, C

  • B. D, B, A, C

  • C. B, C, A, D

  • D. B, C, D, A

Answer: A. D, A, B, C. First compare exponents to learn which mantissa must shift. Next align the binary points, add or subtract the mantissas, and normalise the result. Arithmetic before alignment would combine digits with different place values.

Questions 9 and 10 point to the Pipelining Basics PYQ Questions hub because their source-course pages do not expose stable question-specific titles.

Question 9: what a data interlock does

What is the primary function of data interlocks in a pipelined processor?

  • A. To detect and resolve dependencies between instructions to ensure correct execution

  • B. To increase clock speed by reducing pipeline stages

  • C. To eliminate the need for register

  • D. To disable the pipeline during cache-misses

Answer: A. To detect and resolve dependencies between instructions to ensure correct execution. Suppose I1: R1 = R2 + R3 is followed by I2: R4 = R1 + R5. The interlock detects that I2 needs R1 before I1 has produced it, then stalls or enables the processor's available resolution mechanism. The interlock does not itself always forward data.

Question 10: name the overlap technique

Optimization of each instruction in the processor is done through a technique known as

  • A. RISC

  • B. CISC

  • C. Pipelining

  • D. Execute cycle

Answer: C. Pipelining. RISC and CISC describe instruction-set approaches, while execute is one stage or phase. Pipelining is the technique that overlaps stages of several instructions to improve the aggregate completion rate.

7. Pipelining trap checklist and the next practice step

Before you lock an answer, check five things:

  1. Use the slowest stage to set the clock.

  2. Add the register delay once per clock period.

  3. Use k + n - 1 only for ideal, stall-free operation.

  4. Keep nanoseconds, microseconds and milliseconds consistent.

  5. Separate single-instruction latency from multi-instruction throughput.

Take a 20-second self-test with Question 1. You should be able to say: slowest stage 160 ns, clock 165 ns, 104 cycles, total 17,160 ns, without reopening the solution.

Continue with the GATE Guidance course for the full Computer Architecture sequence and the GATE Test Series for timed practice. Then solve Cache Memory: Mapping and Hit Ratio as your next Computer Architecture problem set.