Performance review findings: encoding and decoding
Two read-only reviews of branch m2-generic-records, done on 2026-09-25:
- Binary primitives:
AvroReader,AvroWriter,PooledBufferWriter. - Generic data model:
GenericDatumReader,GenericDatumWriter,AvroValue,GenericRecord.
The numbers come from scratchpad benchmarks (BenchmarkDotNet and Stopwatch) on an i7-12800H with .NET 10 and TieredPGO. They show direction only. Re-measure with the gated benchmark suite before acting on any of them.
Target frameworks are net10.0;net9.0;net8.0;netstandard2.1;netstandard2.0, and net8+ is the fast path. Unsafe and MemoryMarshal work on every target. Vector128, Bmi2, BitOperations, Utf8.* and IndexOfAnyExceptInRange need net8+.
0. Fix first: security and robustness
These three are confirmed in the code. Status: all three fixed in PR #4 (tests in tests/AvroSharp.Tests/Generic/HostileInputTests.cs).
- A few bytes of input can force a huge allocation.
GenericDatumReader.cs:240setslist.Capacityto the block count before it reads any items.RecordNode.MinimumSizeis0(:189), so everyarray<record>gets the zero-size capMaxZeroSizeItemsPerBlock = 1 << 24(:27). At 16 bytes perAvroValue, a 5-byte input can trigger a 256 MB allocation on the large object heap. A zero-size item such asNullNodeconsumes no bytes, so repeating blocks grow memory without limit.- Compute a real minimum size for a record once its fields are built. A recursive reference counts as 0, so cycles are safe.
- When
MinimumSize == 0, pre-allocate at mostMath.Min(n, 1024)items. - Add a per-read budget for zero-size items.
- Fixed: records compute their minimum size from their fields; arrays and maps pre-allocate at most 1,024 items; zero-size items draw from a per-read budget (
GenericDatumReaderOptions.MaxZeroSizeItems, default 65,536).
- Deeply nested input can crash the process. Nothing limits how deep a recursive schema can nest while reading or writing. A schema like
Node{children: array<Node>}needs about 2 bytes per level, so a crafted input of about 20–50 KB overflows the stack, and a stack overflow cannot be caught. Add an explicit depth counter (default 64–128) and pass it throughRead/Write.- Fixed:
GenericDatumReaderOptions.MaxDepthandGenericDatumWriterOptions.MaxDepth(default 128); the writer limit also stops a record that contains itself.
- Fixed:
- An overflow throws the wrong exception.
list.Count + n(:240) can overflowintacross blocks. It then throwsArgumentOutOfRangeExceptioninstead ofAvroDataException.- Fixed: the running item count is checked against the largest .NET array length before each block, as a
long.
- Fixed: the running item count is checked against the largest .NET array length before each block, as a
1. Binary primitives (AvroReader / AvroWriter)
1.1 Branchless PEXT/PDEP varint paths (biggest win on realistic data)
Status: implemented in the branchless-varints PR, awaiting measurement. One and two bytes are also handled branchlessly (inline), because the mixed 1-2 byte case is the most common and the "1-byte check then PEXT" variant was slower there.
AvroReader.cs:446-653, AvroWriter.cs:247-326. Today a chain of length checks (1, 2, 3/4, 5-8 and 9/10 bytes) handles each varint. That is fast when value lengths are predictable and slow when they are not.
Time per value, in ns, measured over 200,000 values:
| Data | Decode, current | Decode, 1-byte check then PEXT | Decode, fully branchless | Encode, current | Encode, branchless PDEP |
|---|---|---|---|---|---|
| All 1-byte | 0.75 | 0.67 | 2.87 | 0.70 | 1.73 |
| All 3-byte | 2.24 | 3.25 | 2.82 | 1.70 | 1.67 |
| All 8-byte | 3.96 | 2.96 | 2.64 | 3.00 | 1.70 |
| Random 1–8 bytes | 7.9 | 3.64 | 2.79 | 5.8 | 1.73 |
| Random 1–2 bytes | 4.06 | 4.9 | 2.71 | 3.95 | 1.74 |
Recommendation:
- Keep the inline
< 0x80fast path. - Replace the 2 to 8-byte checks with one branchless path, used when
Bmi2.X64.IsSupportedand at least 10 bytes are available. - Fall back to the existing
CompactVarint/SpreadVarint, which is nearly as fast.
// decode
ref var p = ref Unsafe.Add(ref MemoryMarshal.GetReference(_span), (nint)(uint)position);
ulong first = p; if (first < 0x80) { _position = position + 1; return first; }
var word = Unsafe.ReadUnaligned<ulong>(ref p);
var stops = ~word & 0x8080808080808080UL;
if (stops != 0) {
var keep = ((stops & (0UL - stops)) << 1) - 1; // bytes through the terminator
_position = position + (BitOperations.TrailingZeroCount(stops) >> 3) + 1;
return Bmi2.X64.ParallelBitExtract(word & keep, 0x7F7F7F7F7F7F7F7FUL);
}
// 9/10-byte: existing code
// encode (value < 2^56)
// bits = 64 - Lzcnt(value | 1); length = (bits + 6) / 7;
// word = Pdep(value, 0x7F7F..) | (0x8080.. & ((1UL << ((length - 1) * 8)) - 1));
// WriteUnaligned(word); _buffered += length;
- Effect: 2–3× faster on streams of mixed lengths, 1.3–1.5× on 8-byte values such as timestamps, no change on 1-byte values, and about 30% slower on uniform 3-byte values.
- Risk: PEXT and PDEP are microcoded on AMD Zen 1/2 (about 18 cycles), and .NET cannot detect those CPUs. Either accept the cost, or use the shift-and-mask fallback everywhere.
1.2 The benchmarks let the CPU memorize the data
VarintBenchmarks.cs:21,50-56 uses 1,000 values of one length. BinaryEncodingBenchmarks.cs:51-56 repeats i % 20 and BulkReadBenchmarks.cs:34 repeats i % 10. When the same 1,000 values are replayed, the branch predictor learns them, even random ones: random 1–8-byte data costs 2.9 ns/value with 1,000 values and 7.9 ns/value with 200,000.
The PERF comments that chose branchy code (AvroReader.cs:479-485, AvroWriter.cs:277-284) were based on this memorized data.
- Use at least 64K values, or a new random dataset per iteration.
- Add a workload with truly mixed lengths.
- Benchmark through a
ref AvroReaderparameter. A local reader gets promoted to registers, which hides the cost of storing and reloading its position field. - Fixed in the benchmark-fixes PR: VarintBenchmarks, BinaryEncodingBenchmarks and BulkReadBenchmarks use 64K random values with no repeating pattern, VarintBenchmarks adds Mixed1-10 and Mixed1-2, and AvroSharp is called through non-inlined
refhelpers. The PERF comments are marked for re-measurement.
1.3 Decode one-byte runs with SIMD, not just detect them (net8+, Vector128)
AvroReader.cs:322-335, 365-378. OneByteRunLength already uses SIMD to find a run of 1-byte values, but decoding the run is a scalar loop of 16 with two bounds checks per element. Decode all 16 bytes at once instead and store all 16 results unconditionally. This is safe because destination.Length - i >= 16 is already checked, and every slot past the run is overwritten later.
var v = Vector128.Create(chunk);
var z = ((v >> 1) ^ (Vector128<byte>.Zero - (v & Vector128.Create((byte)1)))).AsSByte(); // zigzag, fits sbyte
// widen sbyte -> short -> int (-> long); store 4 x Vector128<int> or 8 x Vector128<long> at destination[i]
_position += run; i += run;
- Effect: about 2–4× faster bulk reads of small values.
- Remaining weakness: data that alternates short and long values pays a 16-byte load and mask per 1-byte value. Masked VByte would fix that but is a large, table-driven change. Defer it until mixed arrays show up in the benchmarks.
- Arm64:
ExtractMostSignificantBitstakes 6–8 instructions there, so re-measure on Arm64.
1.4 Code that stops the JIT from inlining (cheap, helps broadly)
Status: done in the cheap-perf-wins PR (throw helpers, out-of-line float/double/string slow paths, unchecked stores and reads after capacity checks, one slice for 3/4-byte varints).
EnsureRemaining(AvroReader.cs:683-689) andReadLength(:671-681) build interpolated error strings inline, so the JIT does not inline them. Move the throws into[DoesNotReturn][MethodImpl(NoInlining)] statichelpers. Do the same inReadBoolean(:97-102).ReadFloat/ReadDouble(:123-150) containstackalloc, and the JIT never inlines a method that usesstackalloc. Move the multi-segment path into aNoInliningslow method. Thetry/finallyinReadString's pooled path has the same problem.- The fast paths still carry redundant bounds checks:
_buffer[_buffered++]afterEnsure(AvroWriter.cs:90,259,266-267).- The slice plus
BinaryPrimitivesin float/double reads and writes. UseUnsafe.ReadUnaligned/WriteUnalignedguarded byBitConverter.IsLittleEndian. TryReadThreeOrFourByteVarint(:491-503) does 4 bounds checks. Slice once so the constant-index checks fold away.
- Effect: about 1–2 ns per primitive, on all targets.
[SkipLocalsInit]would requireAllowUnsafeBlocksfor a negligible gain, so skip it.
1.5 WriteString in one transcoding pass (net8+)
AvroWriter.cs:139-153 calls GetByteCount and then GetBytes.
- For
value.Length <= 21, the UTF-8 length is at most 63, so the length prefix is always one byte. CallEnsure(1 + 3*len), thenUtf8.FromUtf16(value, buf[1..], ...), then write the prefix. - For longer strings, reserve the worst-case prefix, transcode, and move the bytes down if the prefix turns out shorter.
- Effect: about 15–25% faster short-string encoding.
- The read side (
Encoding.UTF8.GetString) is already optimal.
1.6 Bulk paths that don't exist yet
- Booleans:
ReadBooleans(Span<bool>)validates withIndexOfAnyExceptInRange((byte)0, (byte)1)and then does one memcpy.WriteBooleansisWriteRaw(MemoryMarshal.AsBytes(values)). About 10× faster. WriteLongs/WriteInts: check 8 values at a time with SIMD. If all are in [-64, 63], zigzag and narrow them into one 8-byte store; otherwise use the PDEP path.- Float/double arrays are already a single memcpy on little-endian.
1.7 Buffer management (mostly fine)
Status: the WriteVarintExact point is done (an IBufferWriter destination grows instead); the others remain.
PooledBufferWriterdoubles on growth and returns the whole free tail. That is good.AvroWriter.cs:251-255, 371-381: when fewer than 10 bytes remain, every varint takes the slowWriteVarintExactpath. With anIBufferWriterdestination, grow the buffer and retake the fast path instead.WriteBytes(:123-127) checks capacity twice. For small values, oneEnsure(10 + value.Length)is enough.WriteStringneeds one contiguous span, which contradicts the comment at:415that says large values never need one huge span.
1.8 Already optimal, or not worth doing
- Already optimal: the inline 1-byte path, the
CompactVarint/SpreadVarintbit tricks, float/double arrays as memcpy, zero-copyReadBytesSpan, and the slow paths already split out. - A fully branchless decode with no 1-byte check is 4× slower on all-1-byte streams.
- Vector256/512 and Masked VByte don't pay off for scalar fields, which are decoded one at a time.
- No correctness bugs were found in the primitives. Checked: bounds and lengths, truncated input, overlong varints, the 5th/10th-byte rules, and the invariants of the run path.
2. Generic data model
Already done well:
- Each schema is compiled once into sealed nodes and cached per schema.
- Union writes of primitives are resolved in O(1).
AvroValueis 16 bytes.- Bulk int/long/float/double reads are correct across multiple blocks.
- The in-place fill with
CollectionsMarshal.SetCountis correct.
Measured: dispatch per field costs only about 10% of a read, so turning the node tree into a flat instruction list would not gain much on reads.
2.1 Store primitive arrays as typed arrays (largest win)
GenericDatumReader.cs:232-267, 289-363, GenericDatumWriter.cs:217-230.
- Cost today: decoding 1,000 doubles takes 366 ns, but storing them as
AvroValuebrings the total to 2,917 ns and 16 KB. About 0.65 ns per element is the GC write barrier. Writing takes 2,158 ns, spent on interface indexer calls and virtual dispatch. - Proposal: add a
PrimitiveList<T> : IReadOnlyList<AvroValue>backed by aT[]. It createsAvroValues only when accessed.- The reader decodes blocks straight into the typed array.
- The writer checks for
PrimitiveList<double>and emits one memcpy. - Expose
AvroValue.FromArray(double[])and similar factories so producers get the fast write path too.
- Effect: double arrays read about 6× and write about 7× faster, with half the allocation. Int and long arrays get 2–4×.
- Risk:
AvroValue.cs:25-27promisesList<T>instances. Before 1.0, documentIReadOnlyList<AvroValue>as the only contract.
2.2 Write arrays and maps without interface calls
Status: done in the cheap-perf-wins PR (span writes for List/arrays, struct enumerator for Dictionary, reference checks instead of Kind). The WriteDoubles copy for primitive arrays comes with §2.1.
GenericDatumWriter.cs:219-226. Write from CollectionsMarshal.AsSpan(list) or from AvroValue[] directly. Writing a double array is 1.9× faster this way; copying the values into a double[] and calling WriteDoubles is 2.2× faster.
value.Kind is evaluated twice per array, and each evaluation checks types one by one, ending in an interface cast. Test value.Reference directly instead.
2.3 Union writes of named types do a dictionary lookup per value
Status: done in the cheap-perf-wins PR.
GenericDatumWriter.cs:288-294, 319. Every record, enum or fixed inside a union hashes its full name on every write. Compare against the 1–3 named branches with ReferenceEquals first, and fall back to the name lookup. A [null, Inner] write goes from 54.7 ns to about 13 ns.
2.4 Record field loop
Status: partly done (schema reference check, internal field count, values passed by reference, no try/catch per field; op-codes and ReadInto remain).
- Write (
GenericDatumWriter.cs:177-196): 86 → 66 ns with a flattened loop.- Check
ReferenceEquals(record.Schema, schema)before comparing names. - Add an internal
FieldCountto avoid interfaceCountcalls. - Move the per-field
try/catchoutside the loop. - Then optionally handle primitive fields inline with op-codes.
- Check
- Read (
:191-201): about 10% faster with aReadInto(ref AvroValue slot)shape, which avoids copying through a hidden return buffer. GenericRecord.cs:19makes an interface call to get the field count on every read. Add an internal constructor that takes a cached count.- Fusing runs of fields showed no gain at this layer. Fusion belongs in the source-generated path.
2.5 Allocations
GenericDatumWriter.cs:241:foreachoverIReadOnlyDictionaryboxes the enumerator, costing 64 B per map write. Special-caseDictionary<string, AvroValue>.GenericDatumReader.cs:429: create the map with a capacity ofMath.Min(count, 256)taken from the first block.- Map keys and other repeated strings could use an optional lock-free cache from UTF-8 bytes to
string. Enum symbols already don't allocate. FixedNode.Readallocates abyte[]per fixed. This can't be avoided whileGenericFixedowns abyte[].- A 24-byte
AvroValuelayout would halve the cost of storing an element but use 50% more memory. Not recommended, because §2.1 removes most of that cost.
2.6 Remaining gaps in the bulk paths
- Boolean arrays go through the virtual call per item. Read them with
ReadFixedSpan(n)and validate withVector128. - Double/float bulk reads rent a buffer and then copy twice.
MemoryMarshal.Cast<byte,double>(reader.ReadFixedSpan(n*8))is 20% faster; process 1,024 items at a time.
2.7 Skipping
There is no skip path yet. The writer never emits size-prefixed blocks, although design.md §4.3 says it does.
- Don't size blocks in the generic writer; it would need a pre-pass or double buffering.
- Add
Skip(ref AvroReader)to each reader node for M3 projection:- Arrays of fixed-width items skip
n * widthbytes. - Size-prefixed blocks use their byte size.
- Everything else walks the items.
- Arrays of fixed-width items skip
- Update
design.mdto match.
2.8 JIT notes
Status: done in the cheap-perf-wins PR.
- The primitive writer nodes should compare
ReferenceEquals(value.Reference, PrimitiveMarker.X)instead of callingKind, which checks up to seven types. - In
Kind, check exact types (List<AvroValue>,AvroValue[],Dictionary<string,AvroValue>) before the interface check.
Suggested order
- The security fixes in §0.
- The inlining fixes (§1.4), union matching by identity (§2.3) and span-based array writes (§2.2). These are cheap and safe.
- Fix the benchmarks (§1.2), then re-measure and adopt PEXT/PDEP (§1.1).
- Typed primitive arrays (§2.1).
- SIMD one-byte runs (§1.3), single-pass
WriteString(§1.5), and bulk booleans (§1.6, §2.6). - Skip paths for projection (§2.7), together with M3.