Every clock period has to cover more than just the logic. The registers themselves eat some of it: the launching flip-flop's clock to q delay tcq at the start, and the capturing flip-flop's setup time tsu at the end. Together, tcq + tsu is the register overhead (also called sequencing overhead).
T = tpd(logic) + tcq + tsu
This matters most in pipelining. Splitting logic into N stages divides the logic delay, but each stage still pays the full overhead:
- the clock period can never drop below
tcq + tsu, however thin the stages; - the latency grows by one overhead for each extra stage;
- as stages get thinner, overhead becomes a bigger share of each period, so each extra stage helps less.
That's why real pipelines stop at a moderate number of stages: past a point, the registers cost more than the logic they split.