Skip to content
BetterDL

Pipelining

Also called: pipeline, pipeline stage, pipeline register, pipelined

Splitting a long block of logic into stages separated by registers, so the clock can run faster and several items are in progress at once.

A long block of combinational logic forces a long clock period. Pipelining cuts the logic into stages and puts a register, the pipeline register, between each pair. Each stage now has less logic, so the clock can be faster.

Think of a laundry: wash, dry, fold. Without pipelining, one load finishes all three before the next starts. With pipelining, while load 1 dries, load 2 washes. Each load takes just as long, but loads come out much more often.

Two measures change in opposite ways:

  • Throughput: results per second. With one result per cycle, throughput = f. Pipelining raises it.
  • Latency: time for one item to get all the way through. With N stages it's N × T. Pipelining keeps it the same or makes it longer.

The limits:

  • Each pipeline register adds its own register overhead tcq + tsu to every stage.
  • The clock suits the slowest stage, so stage balancing matters.
  • However thin the stages get, the period can't drop below the register overhead.

The period of an N-stage pipeline is T = (slowest stage's logic delay) + tcq + tsu.

Worked examples

Example

Four stages instead of one

A block of logic has tpd = 24 ns. Registers have tcq + tsu = 1 ns. Compare one stage with four equal stages.

  1. 1.

    One stage: T = 24 + 1 = 25 ns. f = 1000 ÷ 25 = 40 MHz. Latency = 25 ns.

  2. 2.

    Four stages of 6 ns: T = 6 + 1 = 7 ns. f = 1000 ÷ 7 ≈ 142.9 MHz.

  3. 3.

    Latency = 4 × 7 = 28 ns: 3 ns worse, because the overhead is paid four times.

  4. 4.

    Throughput rises from 40 to about 142.9 million results per second, about 3.6 times better.

Example

The ceiling

Registers have tcq + tsu = 40 ps. What's the highest clock frequency any amount of pipelining could reach?

  1. 1.

    Even with almost no logic per stage, each period must cover the overhead: T > 40 ps.

  2. 2.

    f < 1 ÷ 40 ps = 25 GHz.

  3. 3.

    In practice gains shrink well before that, as overhead becomes most of each period.

Common mistakes

  • Thinking pipelining makes each item finish sooner. Latency stays the same or grows; throughput is what improves.

  • Dividing the original period by the number of stages and forgetting the overhead each stage pays.

  • Setting the clock from the average stage. The slowest stage sets it.

Practice Pipelining

Interactive questions with instant feedback and a worked solution for every wrong answer.

Learn it step by step

Pipelining is taught in Timing and Sequential Logic and Basic CPU / Computer Architecture.