A long block of combinational logic forces a long clock period. Pipelining cuts the logic into stages and puts a register, the pipeline register, between each pair. Each stage now has less logic, so the clock can be faster.
Think of a laundry: wash, dry, fold. Without pipelining, one load finishes all three before the next starts. With pipelining, while load 1 dries, load 2 washes. Each load takes just as long, but loads come out much more often.
Two measures change in opposite ways:
- Throughput: results per second. With one result per cycle, throughput = f. Pipelining raises it.
- Latency: time for one item to get all the way through. With N stages it's N × T. Pipelining keeps it the same or makes it longer.
The limits:
- Each pipeline register adds its own register overhead
tcq + tsuto every stage. - The clock suits the slowest stage, so stage balancing matters.
- However thin the stages get, the period can't drop below the register overhead.
The period of an N-stage pipeline is T = (slowest stage's logic delay) + tcq + tsu.