Skip to content
BetterDL

Latency

Also called: pipeline latency

The time one item takes to get all the way through a system, start to finish. In an N-stage pipeline it is N clock periods.

Latency answers "how long does one item take?" It's the time from when an input enters a system until its result comes out.

For a block of logic between registers, the latency is one clock period. For an N-stage pipeline, each item spends one period in each stage:

latency = N × T

Latency is different from throughput, which asks "how many items come out per second?" A car factory might take two days to build one car (latency) but finish one every minute (throughput).

Pipelining usually raises latency a little. The total logic is the same, but each extra pipeline register adds its register overhead, and the period must suit the slowest stage. What pipelining improves is throughput.

Latency appears elsewhere in this topic too: a two flip flop synchronizer adds one or two cycles of latency before the circuit sees an input.

Worked example

Example

Latency before and after pipelining

Logic with tpd = 15 ns sits between registers with tcq + tsu = 1 ns. It's then split into three stages of 4, 5 and 6 ns.

  1. 1.

    Unpipelined: T = 15 + 1 = 16 ns. Latency = 16 ns.

  2. 2.

    Pipelined: the slowest stage sets the period, T = 6 + 1 = 7 ns.

  3. 3.

    Latency = 3 × 7 = 21 ns, 5 ns longer than before.

  4. 4.

    Throughput rose from 1 ÷ 16 ns = 62.5 million to 1 ÷ 7 ns ≈ 142.9 million results per second.

Common mistakes

  • Expecting pipelining to reduce latency. It usually increases it slightly.

  • Adding up the stage delays without overhead. Every stage takes a full period, including tcq + tsu.

  • Mixing up latency (time per item) with throughput (items per time).

Practice Latency

Interactive questions with instant feedback and a worked solution for every wrong answer.

Learn it step by step

Latency is taught in Timing and Sequential Logic and Basic CPU / Computer Architecture.