Pipelining Basics MCQs: 10 Solved Questions with Explanations
Solve ten pipelining questions step by step, including timing calculations, a unit trap, weighted speedup, a non-uniform stage schedule and data interlocks.
KnowledgeGate Team
Exam prep & CS education

Single-instruction latency, steady-state throughput, fill-and-drain time, unequal stages and dependency stalls require different reasoning. Start every numerical by identifying the slowest stage, latch delay, instruction count and units. In GATE preparation, solve each option set before reading the bold answer, and carry every unit through the working.
1. The 60-second pipelining toolkit
For k stages and n independent items with no stalls, total cycles are k + n - 1. The clock period is max(stage delay) + pipeline-register delay, and total time is cycles multiplied by the clock period. These formulas assume a constant clock and no bubbles.
A 5-stage pipeline needs 5 cycles for its first instruction. It then ideally completes one instruction per cycle, so 10 instructions take 5 + 10 - 1 = 14 cycles, not 10 or 50. Pipelining normally improves throughput, not the latency of one instruction.
Three losses matter. Unequal stage delays make every stage follow the slowest stage, dependencies cause data hazards, and shared hardware causes structural hazards. Read Pipelining in Computer Architecture Explained for the complete concept lesson.
2. Questions 1-2: clock period, pipeline fill and total time
Question 1: five stages and 100 independent instructions
A five-stage pipeline has stage delays of 150,120,150,160 and 140 nanoseconds. The registers that are used between the pipeline stages have a delay of 5 nanoseconds each. The total time to execute 100 independent instructions on this pipeline, assuming there are no pipeline stalls, is _______ nanoseconds.
Options: none; this is a numerical-answer question.
Answer: 17160 nanoseconds. The slowest stage takes 160 ns, so the clock period is 160 + 5 = 165 ns. The pipeline needs 5 + 100 - 1 = 104 cycles. Therefore, total time is 104 x 165 = 17,160 ns. Do not sum all five stage delays for every instruction because the instructions overlap in execution.
Question 2: four stages and 1,000 data items
A 4-stage pipeline has the stage delays as 150, 120, 160 and 140 nanoseconds respectively. Registers that are used between the stages have a delay of 5 nanoseconds each. Assuming constant clocking rate, the total time taken to process 1000 data items on this pipeline will be
A. 120.4 microseconds
B. 160.5 microseconds
C. 165.5 microseconds
D. 590.0 microseconds
Answer: C. 165.5 microseconds. Here, Tclk = 160 + 5 = 165 ns, and the cycle count is 4 + 1000 - 1 = 1003. Thus, 1003 x 165 = 165,495 ns = 165.495 microseconds, which rounds to 165.5 microseconds.
3. Questions 3-4: catch a unit mismatch and combine unequal improvements
Question 3: the unit-conversion trap
A 4-stage pipeline has the stage delay as 150,120,160 and 140 ns respectively. Registers that are used between the stages have delay of 5 ns. Assuming constant clocking rate, the total time required to process 1000 data items on this pipeline is
A. 160.5 ms
B. 165.5 ms
C. 120.5 ms
D. 590.5 ms
Answer: none of the listed options. The calculation is 1003 x 165 ns = 165,495 ns, which converts to 0.165495 ms. Option B would match only if the stated delays were microseconds, since 165,495 microseconds = 165.495 ms. Carry the unit through every multiplication before choosing an option.
Question 4: weighted overall speedup
In an enhancement of a design of a CPU, the speed of a floating point unit has been increased by 20% and the speed of a fixed point unit has been increased by 10%. What is the overall speedup achieved if the ratio of the number of floating point operations to the number of fixed point operations is 2:3 and the floating point operation used to take twice the time taken by the fixed point operation in the original design?
A. 1.155
B. 1.185
C. 1.255
D. 1.285
Answer: A. 1.155. Let the original fixed-point time be f, so one floating-point operation takes 2f. The original time is 2(2f) + 3(f) = 7f. The new time is 2(2f/1.2) + 3(f/1.1) = 200f/33. Overall speedup is 7f / (200f/33) = 231/200 = 1.155.
4. Question 5: schedule a non-uniform four-stage pipeline
Question 5: eight jobs across four stages
Consider a 4 stage pipeline processor. The number of cycles needed by the four instructions I1, I2, I3, I4 in stages S1, S2, S3, S4 is shown below:
Instruction | S1 | S2 | S3 | S4 |
|---|---|---|---|---|
I1 | 2 | 1 | 1 | 1 |
I2 | 1 | 3 | 2 | 2 |
I3 | 2 | 1 | 1 | 3 |
I4 | 1 | 2 | 2 | 2 |
What is the number of cycles needed to execute the following loop? For (i=1 to 2) {I1; I2; I3; I4;}
A. 16
B. 23
C. 28
D. 30
Answer: B. 23. Treat I1, I2, I3, I4, I1, I2, I3, I4 as eight jobs moving through four stages. For job j and stage s, use C[j,s] = max(C[j-1,s], C[j,s-1]) + p[j,s]. The stage-4 completion times are 5, 10, 13, 15, 16, 18, 21, 23, so the final instruction completes at cycle 23. Option 16 ignores stage conflicts, while 28 and 30 miss legal overlap.

5. Questions 6-7: why performance falls and how ideal completion works
Question 6: three causes of lost throughput
The performance of a pipelined processor suffers if
A. the pipeline stages have different delays
B. consecutive instructions are dependent on each other
C. the pipeline stages share hardware resources
D. All of the above
Answer: D. All of the above. Unequal delays set the clock by the slowest stage and leave faster stages idle. A dependency can insert a data-hazard stall. Shared hardware can insert a structural-hazard stall.
Question 7: ten instructions in a classic five-stage pipeline
A CPU has a 5-stage pipeline with the following stages Fetch (F), Decode (D), Execute (E), Memory (M) and Write-back (W). Each stage takes one clock cycle to complete. Assume there are no pipeline stalls and the pipeline is initially empty. How many clock cycles are required to complete the execution of 10 instructions?
A. 10
B. 14
C. 15
D. 19
Answer: B. 14. The first instruction completes after 5 cycles, and the other 9 ideally complete one per cycle: 5 + 9 = 14. Equivalently, k + n - 1 = 5 + 10 - 1 = 14. Latency is 5 cycles/instruction, while steady-state throughput is 1 instruction/cycle.
6. Questions 8-10: operation order, interlocks and the core definition
Question 8: arithmetic-pipeline order
The right sequence of suboperations that are performed in arithmetic pipeline is –
A. Align the mantissas
B. Add or subtract the mantissas
C. Normalize the result
D. Compare the exponents
Choose the correct answer from the options given below:
A. D, A, B, C
B. D, B, A, C
C. B, C, A, D
D. B, C, D, A
Answer: A. D, A, B, C. First compare exponents to learn which mantissa must shift. Next align the binary points, add or subtract the mantissas, and normalise the result. Arithmetic before alignment would combine digits with different place values.
Questions 9 and 10 point to the Pipelining Basics PYQ Questions hub because their source-course pages do not expose stable question-specific titles.
Question 9: what a data interlock does
What is the primary function of data interlocks in a pipelined processor?
A. To detect and resolve dependencies between instructions to ensure correct execution
B. To increase clock speed by reducing pipeline stages
C. To eliminate the need for register
D. To disable the pipeline during cache-misses
Answer: A. To detect and resolve dependencies between instructions to ensure correct execution. Suppose I1: R1 = R2 + R3 is followed by I2: R4 = R1 + R5. The interlock detects that I2 needs R1 before I1 has produced it, then stalls or enables the processor's available resolution mechanism. The interlock does not itself always forward data.
Question 10: name the overlap technique
Optimization of each instruction in the processor is done through a technique known as
A. RISC
B. CISC
C. Pipelining
D. Execute cycle
Answer: C. Pipelining. RISC and CISC describe instruction-set approaches, while execute is one stage or phase. Pipelining is the technique that overlaps stages of several instructions to improve the aggregate completion rate.
7. Pipelining trap checklist and the next practice step
Before you lock an answer, check five things:
Use the slowest stage to set the clock.
Add the register delay once per clock period.
Use
k + n - 1only for ideal, stall-free operation.Keep nanoseconds, microseconds and milliseconds consistent.
Separate single-instruction latency from multi-instruction throughput.
Take a 20-second self-test with Question 1. You should be able to say: slowest stage 160 ns, clock 165 ns, 104 cycles, total 17,160 ns, without reopening the solution.
Continue with the GATE Guidance course for the full Computer Architecture sequence and the GATE Test Series for timed practice. Then solve Cache Memory: Mapping and Hit Ratio as your next Computer Architecture problem set.
Keep learning

Instruction Formats and Addressing Modes MCQs: 12 Solved Cross-Concept Questions
Solve 12 cross-concept COA questions that connect addressing-mode choices with PC rules, memory references, opcode fields, and byte-aligned instructions.

Interrupt-Driven I/O MCQs: 12 Solved Questions with Explanations
Attempt 12 previous-year interrupt-driven I/O questions, then check each answer with a concise explanation. The set covers ISR order, vectoring, priority, and CPU-time numericals.

Bitmap and Pixmap MCQs: 12 Solved Pixel Depth and Memory Questions
Build a reliable pixel-memory method through 12 live MCQs covering bitmaps, pixmaps, lookup tables, uncompressed storage, refresh rates, masks, and dithering.

Assembly & Assembler Design MCQs: 11 Solved PYQs Explained
Attempt 11 published PYQs on assembler directives, language levels, tables, register-pair instructions, debugging and fixed-width arithmetic, then check each worked explanation.