For vs while in C, C++ and Go: the assembly is identical

2026-10-02 · Systems programming

"Is for faster than while?" is one of those questions everybody has an opinion on and almost nobody has checked. I write C, C++ and Go, so I checked in all three: compiled the same loop both ways, diffed the machine code, then benchmarked it. The short answer is that the compiler can't tell the two apart. The long answer is that the question sent me to four things that do change loop performance, and to one benchmarking trap that gave me three fake speed differences before I understood it.

The short version

  • for and while produce identical machine code in all 16 C/C++ configurations I tried (gcc and clang, -O0 to -O3) and in every Go pair I tried. The only difference anywhere is one xor scheduled one slot earlier at gcc -O3.
  • A 32-bit unsigned counter against a 64-bit bound made gcc stop vectorizing: 2.4× slower.
  • strlen(s) in the loop condition made a loop quadratic: 835 ms vs 0.19 ms at 320 000 characters.
  • C++ range-for and iterator loops are 2–3× slower than index loops at -O0 and identical at -O2.
  • Go's bounds check, in a simple sum loop, costs nothing I could measure.
  • The trap: the same instructions ran 2× slower when the loop happened to straddle a 64-byte boundary. That produced phantom differences between variants that were byte-for-byte the same loop.

Setup

The kernel is the smallest loop that can't be deleted: sum an array of uint32_t into a uint64_t. Four spellings in C, the same four in C++ and Go.

uint64_t sum_for(const uint32_t *a, size_t n) {
    uint64_t s = 0;
    for (size_t i = 0; i < n; i++) s += a[i];
    return s;
}

uint64_t sum_while(const uint32_t *a, size_t n) {
    uint64_t s = 0;
    size_t i = 0;
    while (i < n) { s += a[i]; i++; }
    return s;
}
// plus sum_dowhile (hand-inverted loop with an n == 0 guard)
// and sum_ptr (walks a pointer to a + n)

Everything ran on a Ryzen 5 4500U laptop (Zen 2), with gcc 13.3, clang 18.1.3 and Go 1.27.1, pinned to one core with taskset. Each timing is a median: 15 samples per variant, interleaved round-robin after a warm-up round so clock drift hits every variant equally, 228 elements per sample, and the whole thing repeated five times. The array is 4096 elements (16 KB) so it stays in L1 and the loop is the only thing being measured. The laptop boosts, so absolute nanoseconds move between runs: the same scalar loop took 0.26 ns, 0.28 ns, and 0.42 ns when the clock dropped to its 2.4 GHz base. Within a run, the ratios are stable, so those are what I trust.

Test 1: is the machine code the same?

Timing a loop tells you whether two versions behave differently. Diffing the code tells you whether they are different at all, which is a stronger and cheaper test, so that came first. Compile to an object file, disassemble each function, strip the addresses and jump targets (they differ because the functions sit at different offsets), and compare:

body() {
  objdump -d --no-show-raw-insn "$1" |
  awk -v f="<$2>:" '$2==f{p=1;next} /^$/{p=0} p' |
  sed -E 's/^\s*[0-9a-f]+:\s*//; s/\s*#.*//; s/<[^>]*>//g;
          s/^(j[a-z]+|call)\s+.*/\1 T/; s/\s+/ /g; s/ $//' |
  grep -Ev '^(nop|cs nop|data16)'
}
diff <(body loops.o sum_for) <(body loops.o sum_while) && echo IDENTICAL

My first version of this script reported every pair as different. The jump targets were the problem: je 20 in one function and je 50 in the other are the same jump. Normalizing them fixed it. = means byte-for-byte identical after normalization, ≈ means the same instructions in a different order:

-O0-O1-O2-O3
gcc (C)===≈
clang (C)====
g++ (C++)===≈
clang++ (C++)====

The single ≈ is gcc -O3, where one xorl %edx, %edx is placed one slot earlier in sum_for. That is the scheduler, not the loop. And it makes sense that the table is so uniform: both front ends lower for and while to the same control-flow graph, a header block with the test, a body, and a back edge. There is no for in the IR to be fast or slow. This is what gcc -O2 produced for both (without the leading endbr64); it already turned the indexed loop into a pointer bump:

    testq   %rsi, %rsi
    je      .Lzero
    leaq    (%rdi,%rsi,4), %rcx     # end = a + n
    xorl    %eax, %eax
.Lloop:
    movl    (%rdi), %edx
    addq    $4, %rdi
    addq    %rdx, %rax
    cmpq    %rcx, %rdi
    jne     .Lloop
    ret
.Lzero:
    xorl    %eax, %eax
    ret

Go gives the same answer. With go tool objdump, for i := 0; i < len(a); i++, for i < len(a) with a manual i++, for i := range a and for _, v := range a are four identical 11-instruction functions. Even with optimizations off (-gcflags='all=-N -l'), for and while keep the same 35 instructions; one JMP sits in a different place. (Note the all=: without it, -gcflags only applies to the package named on the command line. My first -N run silently compiled the loops with optimizations on and "proved" everything was identical.)

    XORL CX, CX                 // sum
    XORL DX, DX                 // i
    JMP  cond
body:
    MOVL 0(AX)(DX*4), SI
    ADDQ SI, CX
    INCQ DX
cond:
    CMPQ BX, DX
    JG   body
    MOVQ CX, AX
    RET

The hand-written do-while is the one variant that is not identical. Writing if (n == 0) return 0; yourself is doing by hand what the compiler's loop inversion already does for a for, so the surrounding code differs slightly, but the hot loop is the same.

Test 2: what does change a loop

If the keyword is irrelevant, the interesting question is what isn't. Four things came up.

The counter type decides whether the loop vectorizes

int i against an int n and size_t i against a size_t n both vectorize. unsigned i against a size_t n does not, in gcc:

uint64_t sum_u32(const uint32_t *a, size_t n) {
    uint64_t s = 0;
    for (unsigned i = 0; i < n; i++) s += a[i];   // 32-bit counter, 64-bit bound
    return s;
}

If n is larger than UINT_MAX, i wraps to zero and the loop never ends, so the compiler has to preserve that. Signed overflow is undefined behavior, so the compiler may assume an int counter never wraps, which is why int is fine. gcc -O3 compiled the unsigned version to a 17-instruction scalar function with no SIMD instructions at all. clang vectorizes it anyway, with a much longer function (50 instructions vs 17, presumably a runtime check plus a scalar fallback; I didn't read it closely). Times in ns per element:

gcc -O2gcc -O3gcc -O3, alignedclang -O3
for0.2570.1160.1170.111
while0.2560.1170.1420.111
do-while0.2570.1420.1160.111
pointer0.2570.1160.1420.141
int counter0.2570.1420.1160.116
unsigned counter0.2580.2810.5550.111

Read the gcc -O3 column for the real effect: vectorizing takes the loop from 0.257 (gcc -O2 is scalar, one element per cycle) to 0.116 ns, and the unsigned counter loses it, landing at 0.281. The scattered 0.142 values in the gcc -O3 columns are not a finding. They are the trap from the last section, and they move when I change where the loops sit in memory. Ignore them for now.

The takeaway is a habit, not a benchmark result: use size_t or a signed type for the counter, and never a narrower unsigned type than the bound it's compared against.

A function call in the condition runs every iteration

The loop condition is evaluated on every iteration whether it's spelled for or while. Put a call there and you pay for it every time, unless the compiler can prove the result doesn't change:

void upcase_call(char *s) {
    for (size_t i = 0; i < strlen(s); i++)
        if (s[i] >= 'a' && s[i] <= 'z') s[i] -= 32;
}

void upcase_hoist(char *s) {
    size_t n = strlen(s);
    for (size_t i = 0; i < n; i++)
        if (s[i] >= 'a' && s[i] <= 'z') s[i] -= 32;
}

The store to s[i] is a char write, and char may alias anything, including the string's own terminator. So the compiler can't prove strlen returns the same value next iteration. Both gcc and clang at -O2 keep the call inside the loop, and the timing shows it:

Lengthstrlen in conditionhoistedRatio
10 0000.437 ms0.0052 ms84×
40 00014.4 ms0.0214 ms674×
160 000204 ms0.0934 ms2 189×
320 000836 ms0.190 ms4 402×

gcc -O2; clang is within a few percent (820 ms at 320 000). Doubling the length quadruples the time, plus a jump between 20 000 and 40 000 characters where the string stops fitting in L1. Go doesn't have this exact trap, because len(s) is a field in the slice header, but the rule is the same in all three languages: i < f() calls f on every iteration.

C++ at -O0 is where the folklore comes from

With std::vector, the index loop, the range-for, and the iterator loop all do the same work. At -O0, they don't compile to the same thing, because every operator[], operator*, operator++ and operator!= is a real function call. Ratios to the index for loop (4.59 ns/elem at -O0, 0.258 at -O2):

-O0-O2-O2, aligned
index for1.00×1.00×1.00×
index while1.00×0.98×1.00×
range-for2.18×0.99×1.00×
iterator while3.09×1.00×1.00×
std::accumulate2.09×1.96×0.99×

The -O0 column explains why people who benchmark a debug build conclude "range-for is slower". The index for and while are still byte-identical there, so the keyword is not the cause; the abstraction layer is. At -O2 they all collapse to the same loop, except one number: std::accumulate at 1.96×. That one is not real, and the "aligned" column is the hint. It's the subject of the last section.

Go: bounds checks, and the one place for and while differ

Go checks every a[i] against the slice length unless the compiler can prove it's in range. -d=ssa/check_bce/debug=1 shows which accesses kept their check:

// i is bounded by len(a): the check is proved away
for i := 0; i < len(a); i++ { s += uint64(a[i]) }
for i := range a             { s += uint64(a[i]) }

// i is bounded by an unrelated n: Found IsInBounds, checked every iteration
for i := 0; i < n; i++ { s += uint64(a[i]) }

// re-slice once up front: the check moves out of the loop (Found IsSliceInBounds)
a = a[:n]
for i := 0; i < n; i++ { s += uint64(a[i]) }

Both the for i < n and the while i < n forms keep the check, and the a = a[:n] hint removes it from both, so the loop keyword is irrelevant here too. The machine code confirms it: the checked loop has an extra CMPQ/JA pair that jumps to runtime.panicBounds. What does that pair cost? To isolate it from everything else I wrote the same loop by hand in assembly, with and without a per-iteration compare-and-branch that is never taken, both inside one cache line. Medians of three runs were 0.2815 and 0.2819 ns per element, and the minimums were identical at 0.257. In a sum loop, where the critical path is the add chain and the predicted branch fits in spare issue slots, a bounds check was not measurable. That is a result about this loop on this CPU; a check can still matter when it stops a larger optimization, and Go's compiler doesn't auto-vectorize, so here it isn't blocking one.

The one real semantic difference between for and while in Go is the loop variable. Since Go 1.22, a variable declared in the for header is a fresh copy per iteration. A while-style loop with the counter declared outside is one shared variable:

for i := 0; i < 3; i++ { fs = append(fs, func() int { return i }) }
// calling each closure: [0 1 2]

i := 0
for i < 3 { fs = append(fs, func() int { return i }); i++ }
// calling each closure: [3 3 3]

I ran both to be sure. The per-iteration copy only costs anything when a closure captures the variable; a for loop that captures nothing allocates nothing (0 allocations per run). If you have ever been bitten by the old "all closures see the last value" bug, for with the variable in the header is now the safe spelling.

The trap: the same loop, 2× slower

Everything above is clean because I'd learned to distrust the timings by then. Three times, the benchmark reported a large difference between variants whose hot loops were instruction-for-instruction the same.

1. gcc -O3 do-while, 22% slower. The do-while variant measured 0.142 against 0.116 ns/elem for for, reproducibly, across runs. The obvious conclusion is "hand-inverting the loop made it worse". But the inner vector loop in both functions is the same nine instructions (movdqu, addq, movdqa, two punpck, two paddq, cmpq, jne). The loops sit at different addresses, so I rebuilt with -falign-loops=32. The penalty did not disappear: it moved. Now do-while was fast and while and the pointer variant were the slow ones. That is the "gcc -O3, aligned" column in the table above, and it is also why the int counter row jumps around.

2. std::accumulate, 2× slower. At -O2, std::accumulate ran at about twice the time of the range-for, in every run (0.556, 0.555, 0.507 against 0.28, 0.27, 0.26). Both compile to the same five-instruction, 14-byte loop (load, two adds, compare, branch). In the linked binary the range-for's loop is at 0x1718–0x1725, inside one 64-byte line. accumulate's is at 0x1778–0x1785, and it runs across the boundary at 0x1780. Forcing loop alignment made every variant equal (0.99×, in the table).

3. Go, two groups. Eight Go variants fell into two groups, with nothing in between:

VariantCheckCrosses linens/elem
for i < len(a)nono0.258
while i < len(a)nono0.259
range indexnono0.261
range valuenono0.259
for i < nyesyes0.511
while i < nyesyes0.546
for i < n + hintnoyes0.526
range n + hintnoyes0.533

A first reading: "the bounds check costs 2×". It doesn't survive the last two rows ("hint" is a = a[:n]). Those have no check in the loop and are just as slow. My second suspicion was a NOPL that Go's assembler had inserted inside the slow loops, between INCQ and CMPQ. I tested it in hand-written assembly, with and without the nop: no difference. Reordering the functions in the source file, to see whether the slow group followed the layout, didn't change the grouping either. What did separate the groups was where the hot loop ended up. I extracted the address range of all eight loops from the final binary. The four fast ones (14 bytes each) are every one inside a single 64-byte line, for example 0x49bb6b–0x49bb78. The four slow ones (17 and 22 bytes) all cross a line boundary, for example 0x49bc74–0x49bc84, across 0x49bc80. Eight for eight, and the "Crosses line" column is the only one that predicts the timing.

To prove it, I wrote the same five-instruction loop in assembly and placed it by hand:

Same loop, hand-placedns/elem
inside one 64-byte line0.258
starts 8 bytes before a 64-byte boundary0.508
starts 12 bytes before a 64-byte boundary0.508

Those are the same two numbers the Go variants split into, to three digits: 0.258 and 0.508. 0.25 ns is one cycle at about 3.9 GHz (the clock a one-cycle-per-element scalar loop implies; at the 2.4 GHz base I measured 0.42 ns for the same loop), so a loop that straddles a 64-byte boundary here costs one extra cycle per iteration. I established that placement alone reproduces it; I did not dig into the front-end mechanics of why Zen 2 pays that cycle, and other CPUs will behave differently. The point survives either way.

The first column of the C table, scalar gcc -O2, never showed any of this. Every variant was 0.257. Its loops happened to all land well. Which is the actual problem: whether you see a phantom difference depends on luck.

With the working set at 64 MB instead of 16 KB, the DRAM-resident run, the loop-form differences and the alignment artifacts vanish into memory bandwidth. Every gcc -O3 and clang variant with a signed or 64-bit counter runs at 0.28 ns per element (about 14 GB/s), within 2% of each other, whether the loop is for, while, do-while or pointer, aligned or not. What stays visible is vectorization: scalar gcc -O2 can't saturate the bus (0.46 ns), and the unsigned counter loop that gcc leaves scalar is still 1.8× slower at -O3 (0.50 ns). Clang is unaffected (0.28).

What I'd take from this

  • Pick the loop for what it says. for when you have a counter, while when the exit condition is the point, range in Go. There is no performance reason to prefer one.
  • Check these instead: the counter's type against its bound, calls in the condition, and (in Go) whether the compiler can prove the index in range.
  • Never benchmark at -O0. It measures call overhead for abstractions that don't exist in a release build.
  • Diff the assembly before trusting a microbenchmark. If the code is the same and the time differs, suspect placement before suspecting the compiler. -falign-loops=32, interleaved samples, and a second build with the functions in a different order are cheap ways to find out.

All of this is specific to one Zen 2 laptop and these toolchain versions. The identical-assembly result should travel; the alignment numbers should not.