Hybrid varints (825f114): i7-12800H and EPYC 7543
Measures commit 825f114 ("Varints: restore branchy 1-4 byte paths, keep PEXT/PDEP word path for 5-10 bytes"). This commit follows up on 2026-09-26-branchless-varints.md. The baseline is ad53209 / 8e4a4a7, from 2026-09-26-benchmarks-64k.md. Analyzed on 2026-09-26.
| i7 run | EPYC run | |
|---|---|---|
| CPU | Intel Core i7-12800H (Alder Lake) | AMD EPYC 7543 (Zen 3) |
| Code | 825f114 on branch branchless-varints |
825f114 (from the version string in the gate output) |
| Job | DefaultJob | DefaultJob |
| Benchmarks | VarintBenchmarks (64K random values), GenericRecordBenchmarks |
same |
The EPYC comparison against Apache passed 18 of 18.
Per value: baseline → 825f114
Times are per value (÷ 65,536). The numbers in brackets are the fully branchless commit 57c987c.
| Case | i7 | EPYC | Verdict |
|---|---|---|---|
| Decode 1 / 2 bytes | 0.61 / 0.89 → 0.61 / 0.89 ns | 0.81 / 1.08 → 0.82 / 1.08 ns | Back to baseline [2.7–2.8] |
| Decode 3 / 4 bytes | 2.11 / 2.12 → 2.15 / 2.10 | 2.70 / 3.24 → 2.71 / 3.24 | Unchanged |
| Decode 5 bytes | 4.39 → 3.92 | 5.25 → 4.62 | 11–12% faster |
| Decode 8 bytes | 3.95 → 3.57 | 4.88 → 4.87 | 10% faster on the i7 |
| Decode mixed 1–10 | 9.20 → 8.19 | 10.18 → 8.56 | 11–16% faster [6.6–6.8] |
| Decode mixed 1–2 | 3.79 → 3.79 | 3.73 → 3.73 | Unchanged [2.7–2.8] |
| Encode 1 / 2 bytes | 0.77 / 0.92 → 0.76 / 0.93 | 1.37 / 1.35 → 1.37 / 1.34 | Baseline; the EPYC comparison passes again (1.36×) |
| Encode 3 / 4 bytes | 2.51 / 2.72 → 2.45 / 2.56 | 3.79 / 3.80 → 3.78 / 3.80 | Unchanged |
| Encode 5 / 8 bytes | 3.42 / 3.46 → 2.88 / 2.85 | 4.89 / 4.88 → 4.35 / 4.34 | 11–18% faster |
| Encode mixed 1–10 | 7.29 → 6.32 | 8.55 → 7.01 | 13–18% faster [4.9–6.0] |
| Encode mixed 1–2 | 3.78 → 3.61 | 3.72 → 4.24 | 14% slower on the EPYC [1.2–2.5] |
Whole records (GenericRecord)
| i7 before → after | EPYC before → after | vs Apache (i7 / EPYC) | |
|---|---|---|---|
| Read | 920 → 801 ns (13% faster) | 1,046 → 1,000 ns | 2.0× / 2.6× |
| Write | 339 → 335 ns | 540 → 526 ns | 4.0× / 4.4× |
i5-3570K (no BMI2, shift-and-mask fallback)
Same commit, DefaultJob. "Before" is the ad53209 64K run on the same machine. The comparison against Apache passed on every row.
| Case | Before → 825f114 |
vs Apache |
|---|---|---|
| Decode 1 / 2 bytes | 1.85 / 2.06 → 1.85 / 2.06 ns | 0.43 / 0.31 |
| Decode 3 / 4 bytes | 3.62 / – → 4.28 / 4.11 | 0.47 / 0.35 |
| Decode 5 / 8 bytes | 7.41 / 7.64 (no i5 baseline) | 0.53 / 0.37 |
| Decode mixed 1–10 | 12.92 → 11.16 (14% faster) | 0.52 |
| Decode mixed 1–2 | 4.93 → 4.93 | 0.52 |
| Encode 1 / 2 bytes | 2.04 / 2.13 → 2.03 / 2.13 | 0.83 / 0.47 |
| Encode 3 / 4 bytes | 4.89 / – → 4.77 / 5.30 | 0.70 / 0.60 |
| Encode 5 / 8 bytes | 7.62 / 7.72 (no i5 baseline) | 0.70 / 0.45 |
| Encode mixed 1–10 | 10.77 → 9.89 (8% faster) | 0.59 |
| Encode mixed 1–2 | 4.62 → 4.62 | 0.64 |
GenericRecord: read 1,474 ns (0.43× Apache, 0.68× the allocation), write 678 ns (0.22×, no allocation). There is no earlier i5 GenericRecord run on the 64K benchmarks to compare with.
- The i5 has no mixed 1–2 encode regression, and neither does the i7. That points at the EPYC comparison rather than the code (see below).
- 3-byte decode reads 18% slower than the
ad53209run, and slower than 4 bytes (4.28 vs 4.11 ns). The 3- and 4-byte code is the same as in8e4a4a7, so this is code layout or noise in one of the two runs. It is still 0.47× Apache. Unverified.
On the EPYC mixed 1–2 encode regression (finding 3): the EPYC baseline is the earlier 64K run, whose code was unknown and probably predates ad53209 (see 2026-09-26-benchmarks-64k.md). The one- and two-byte write path in 825f114 is the same source as in 8e4a4a7. So the 3.72 → 4.24 ns change may come from ad53209 or from run-to-run variation, not from this commit. An A/B run of 8e4a4a7 against 825f114 on the EPYC, limited to Mixed1-2, would settle it.
Findings
- This is a good trade overall.
- The common short values (1–4 bytes) are back to full speed.
- The PEXT/PDEP path gives a clean 10–18% on long values (5–10 bytes, such as timestamps and ids) on both CPUs.
- Whole records got faster on both machines, and that is the benchmark that matters most.
- The odd 5-byte result is fixed on the EPYC, where 5 bytes now decodes faster than 8 (4.62 vs 4.87 ns). On the i7, 5 bytes is still slightly slower than 8 (3.92 vs 3.57 ns).
- There is one regression: mixed 1–2 byte encode on the EPYC went from 3.72 to 4.24 ns, only 1.33× faster than Apache. The i7 improved slightly on the same case.
- The commit was meant to leave the 1–2 byte path as it was. The likely suspect is a change in code layout or inlining on Zen 3.
- Compare the JIT's generated code (
DOTNET_JitDisasm=WriteVarint*) for8e4a4a7and825f114on the EPYC.
- 1-byte encode on the EPYC is still slow, and this problem is specific to Zen 3. It takes 1.37 ns, 1.8× the i7, and is only 1.36× faster than Apache. The problem predates all the varint work (see 2026-09-26-benchmarks-64k.md, finding 2).
- Mixed-length streams give up some of the branchless gain. Mixed 1–10 is 8.2–8.6 ns now against 6.6–6.8 ns branchless. That is the accepted price of fast 1-byte values.
- The 3–4 byte range is untouched and is still the biggest step up in cost: 2.1 ns on the i7 and 2.7–3.2 ns on the EPYC, against about 1 ns for 2 bytes. Sending 3–4 bytes through the PEXT path might help, but the branchless run made 3 bytes 1.4–1.7× slower. It needs its own measurement.
Next steps
- Find the cause of the EPYC mixed 1–2 byte encode regression (finding 3) before merging
branchless-varints. - Profile EPYC 1-byte encode (finding 4). It is the smallest lead over Apache left in the varint benchmarks.
- Try 3–4 byte values through the PEXT word path, and judge it on
GenericRecordBenchmarksas well as the varint benchmarks (finding 6).