A ripple carry adder is slow because each carry waits for the one below it. But each carry is just a Boolean function of the inputs, so a carry-lookahead adder (CLA) computes all of them directly, in parallel.
Start from the generate and propagate recurrence Cᵢ₊₁ = Gᵢ + PᵢCᵢ, and keep substituting the carry below until only G's, P's and C0 remain:
- C1 =
- C2 =
- C3 =
- C4 =
Read each product term as a story: a carry starts somewhere (at a G, or at C0), and every position above it propagates it.
Speed. In the course's gate delay model, with gates as wide as needed: G and P are ready at 1, every carry (a two-level AND-OR) at 3, and every sum Sᵢ = Pᵢ ⊕ Cᵢ at 4, for any width. A 4-bit ripple adder needs 9 for C4.
Cost. Cᵢ has i + 1 product terms, the longest with i + 1 inputs. At 64 bits that would mean a 65-input AND gate, and building one from small gates adds the delay back. Real designs use small lookahead blocks, typically 4 bits, and repeat the G/P idea between blocks (block carry lookahead).