Changelog

All notable changes to this project are documented here. The format follows Keep a Changelog and the project uses Semantic Versioning.

Unreleased

Breaking

The public API review before 1.0 (#134) renames and moves members, so the API can be frozen without them (#73). Regenerate checked-in generated code with the new avrosharp gen: code from 0.2 calls support members that moved. From 1.0, code generated by a version keeps compiling against later runtimes of the same major version (the policy).

  • Generated-code support: AvroGeneratedCode, AvroRecordPlan, AvroPlanCache, AvroConversion and AvroUninitialized move to AvroSharp.Serialization.Generated.
    • IAvroCodec<T> is now IAvroValueSerializer<T>, the primitive codecs Avro…Serializer, and the generated nested AvroCodec struct ValueSerializer: "codec" means block compression only.
    • Removed, because generated code no longer calls them: GetRecordPlan(writer, reader) without the cache, PutTypeMismatch(object?, string, string), AvroRecordPlan.Target(int) and Conversion(int), and IsSameSchema (use AvroSchema.HasSameCanonicalForm).
  • Exceptions: AvroDataException (was AvroSharp.IO) and AvroSchemaException (was AvroSharp.Schemas) are in AvroSharp.
  • Schema lookup: IAvroSchemaStore is now IAvroSchemaResolver, with GetSchemaAsync, which an implementer has to add. The AvroMessageReader factories name its parameter resolver.
  • AvroSerializer: Serialize(output, value) and TrySerialize(destination, value, out bytesWritten) take the output first.
  • Renamed:
    • GenericDatumReader.Schema and GenericDatumJsonReader.Schema are now WriterSchema;
    • EnumSchema.Default is now DefaultSymbol;
    • AvroReader.Skip(long) is now SkipRaw;
    • AvroValue.FromByteArray, FromGenericRecord and FromGenericFixed are now FromBytes, FromRecord and FromFixed;
    • AvroFileReader<T>.GetMetadataString is now TryGetMetadataString(key, out value);
    • CodeGenOptions.DefaultNamespace, NamespaceMapping and PropertyNaming are now Namespace, NamespaceMap and PropertyNames.
  • Changed types:
    • AvroFileReader<T>.Codec is the AvroCodec (its name is Codec.Name);
    • AvroFileWriterOptions.Metadata holds ReadOnlyMemory<byte> values;
    • SchemaFingerprint.Crc64AvroEmpty is a long;
    • ConfluentSchemaIdHeader.Encode takes an AvroSchemaId.
  • AvroCodec.CreateDeflate is replaced by new DeflateCodec(level).
  • Validation: the generic and schema parse options reject a MaxDepth below 1 and a negative MaxZeroSizeItems. CodeGenOptions rejects nullable annotations with a LanguageVersion below 8, and any version below 7.

Added

  • AvroMessageReader<T>.ReadAsync, which fetches an unknown fingerprint through the resolver once, and AvroSchemaStore.GetSchemaAsync (#134).
  • AvroSchema.HasSameCanonicalForm: whether two schemas have the same encoding. Equals stays reference equality, as AvroSchema's docs now say (#134).
  • DeflateCodec: the built-in codec is public, with Default and Level, as the Codecs package's codecs have (#134).
  • New overloads:
    • AvroSerializer.Deserialize<T>(in ReadOnlySequence<byte>, AvroSchema writerSchema);
    • GenericDatumReader.Read(in ReadOnlySequence<byte>);
    • AvroMessage.Write(output, in AvroValue, GenericDatumWriter);
    • AvroRegistryMessage.Write(output, framing, id, in AvroValue, GenericDatumWriter) (#134).
  • Default limits as constants: GenericDatumReaderOptions.DefaultMaxDepth and DefaultMaxZeroSizeItems, GenericDatumWriterOptions.DefaultMaxDepth and AvroSchemaParseOptions.DefaultMaxDepth (#134).
  • Generator settings (#134):
    • The AvroSharpNamespaceMap MSBuild property, avro.ns:CSharp.Ns entries separated by ;, like the CLI's --namespace-map.
    • Generator properties accept their values in any case.
    • A value the generator doesn't recognize is warning AVROGEN006; before, a typo was ignored silently.
    • --language-version 7 implies --no-nullable.
  • Records built from another record's fields: a record given a field that belongs to another record takes a copy, so new RecordSchema(name, other.Fields.Append(field)) works. It threw before (#134).
  • A dev container (.devcontainer/) and build/ci-local.sh, which build and test as the Linux CI job does. The container has Ubuntu 24.04 with the .NET 8, 9 and 10 SDKs, the Native AOT prerequisites, CI's variables and a cached NuGet volume; the script runs the job's steps in order. DocFX and ReportGenerator are pinned local tools (.config/dotnet-tools.json), used by CI, the docs workflow and the container alike (#139).

Fixed

  • A fixed type named Equals or GetHashCode generated code that did not compile (CS0542); it is renamed, with a note, like the other generated member names (#141).
  • A decimal default that the generated C# decimal cannot hold (beyond 96 bits, or with more digits than the precision) generated code whose constructor threw; it is now a generation error (#141).

Changed

  • Tests for the remaining gaps of the test review (#142): enum and fixed aliases, the resolving reader's options, the JSON writer's widening, depth limit and shape errors, maps of any IReadOnlyDictionary, hostile container files (overlong varints, truncation on the async path, trailing bytes when pipelined), AvroRegistryMessageReader.ReadAsync with a missing schema or a cancelled fetch, logical-value range errors, and control characters in schemas.
  • The nightly fuzzing workflow also runs a random-schema code-generation test (#141): random schemas, with hostile names and every kind of default, are generated, compiled for C# 7.3, 12 and the latest version, and round-tripped. PR CI runs it on 100 schemas.
  • The API reference on the documentation site is built from the net10.0 build, so it shows the .NET 8+ API (AvroSerializer, IAvroSerializable<T> and the overloads that take no delegates), and lists those members with their equivalents on other targets (#136).

0.2.0 - 2026-09-29

The avrosharp command-line tool; generated types that read older versions of their schema, write and read memory without allocating, and start with their schema defaults; and the fixes from an independent review: bounded memory and nesting for hostile input, schema resolution that follows the specification and Java, and code generation fixes. The source generator needs the .NET 10 SDK or Visual Studio 2026 and later; the generated code and the tool run on .NET 8 and later (and the code on .NET Standard 2.0).

Added

  • The avrosharp command-line tool, package AvroSharp.Tool: a dotnet tool like Apache.Avro's avrogen, for .NET 8 and later (#33).
    • avrosharp gen writes the source generator's code for .avsc files, or folders of them, with the generator's options: one file per type, in folders for its namespace, or with --flat in one folder.
    • avrosharp schema canonical and schema fingerprint print the Parsing Canonical Form, and the CRC-64-AVRO, MD5 or SHA-256 fingerprint as hex, base64 or (CRC-64) the decimal Java prints.
    • Files may refer to each other's types in any order. Errors are in the compiler's format, with the generator's diagnostic IDs. The exit code is 0 on success, 1 when the command fails and 2 for an invalid command line.
  • CodeGenOptions.NamespaceMapping (--namespace-map in the tool): C# namespaces for Avro namespaces and the namespaces under them, as avrogen's --namespace a:b. Only the generated C# namespace changes, not the schema. GeneratedSource.Namespace gives each file's C# namespace.
  • SchemaFileSet in AvroSharp.CodeGen: parses schema files that refer to each other's named types, in any order, as the source generator and the tool do.
  • Bulk booleans (#26): AvroReader.ReadBooleans checks and copies a block of booleans as one span, and AvroWriter.WriteBooleans writes one as a single copy. Generated code uses them for boolean arrays, and the generic reader and writer use them for boolean arrays.
  • Generated records implement new interfaces (#115): IAvroWritable (WriteTo(ref AvroWriter)) and IAvroReadable (ReadFrom(ref AvroReader), which fills an existing instance and reuses its lists, dictionaries and records), and on .NET 8 and later IAvroSerializable<T> with the static Schema, Write and Read members. With it, AvroSerializer.Serialize/TrySerialize/Deserialize, AvroFileWriter.Create<T>(stream), AvroFileReader.Open<T>(stream)/OpenAsync<T>, AvroStreamWriter.Create<T>, AvroStreamReader.Open<T>, AvroMessage.ToArray<T>/Write<T>, AvroMessageReader.Create<T>, AvroRegistryMessage.ToArray<T>/Write<T> and AvroRegistryMessageReader.Create<T> need no delegates. Readers of files and messages resolve other versions of the type's schema.
  • Generated records write and read memory without allocating (#115): TryWriteAvroBytes(Span<byte>, out int) (no exception when the value does not fit), WriteAvroBytes(Span<byte>), WriteAvroBytes(IBufferWriter<byte>), FromAvroBytes(data, out int bytesConsumed) and FromAvroBytes(in ReadOnlySequence<byte>).
  • AvroWriter.WriteInts and WriteLongs write array items in bulk.
  • CodeGenOptions.LanguageVersion: the C# version generated code may use. The source generator sets it from the project; with C# 11 or later, generated code adds IAvroSerializable<T> for .NET 8 and later.
  • Generated code is easier to debug and read (#117): records have a [DebuggerDisplay] with their first three simple fields, the serializers are [DebuggerNonUserCode], and a property whose logical type keeps its raw type (a decimal wider than 28 digits, duration, timestamp-nanos, or any logical type with AvroSharpLogicalTypes=raw) says in its documentation what the value means.
  • Fixed types have ==, != and AsSpan() (#117).
  • The generator reports a property renamed to avoid a clash (for example user_id and userId in one record) as informational diagnostic AVROGEN005, naming both fields; GeneratedSource.Notes carries these notes for other callers of CSharpCodeGenerator (#117).
  • Generated files suppress CS8981, so all-lower-case Avro type names compile without warnings (#117).
  • AvroFileReaderOptions.MaxSchemaLength and MaxZeroSizeValuesPerBlock (#129, #130).
  • GenericDatumJsonReader.ReadDefault: converts a schema default to an AvroValue, the value a reader gives a field that the data lacks (#131).

Changed

  • Packages list only the dependencies they use (#20). AvroSharp.Generators and AvroSharp.CodeGen depend only on AvroSharp, and AvroSharp.Codecs only on AvroSharp and its compression libraries; before, they also listed System.Memory, System.Text.Json and other packages that AvroSharp brings. AvroSharp's netstandard2.1 build no longer depends on System.Memory or Microsoft.Bcl.AsyncInterfaces, which that target has built in. The generator still needs the .NET 10 SDK or Visual Studio 2026 and later; docs/design.md records what fails on older SDKs.
  • Faster varint encode (#102). Where PDEP is fast (Intel from Haswell, AMD from Zen 3), 3- to 8-byte values are one inline 8-byte store; elsewhere one out-of-line call writes 3 to 5 bytes with one 4-byte store. On the i7-12800H, 3- to 8-byte encodes are 28–38% faster and 3-byte encode is 1.52× faster than Apache.Avro (1.09× before); on the EPYC 7543 they are 25–35% faster, and on an i5-3570K and a Ryzen 5 3500U without fast PDEP, 3 to 5 bytes are 7–33% faster. 1-byte encode is 6–8% slower on the fast-PDEP machines. Record writes are unchanged or faster; generated writes are 7–8% faster on the i7 and the EPYC.
  • AvroWriter.WriteLongs/WriteInts keep the write position in a local while the buffer has room, instead of storing and reloading it after every value (#27, the per-value round trip through memory that limited one-at-a-time writes on the EPYC). They are at least as fast as single writes for every length on the machines measured, and 1-byte bulk writes are 19–48% faster than single ones.
  • Bulk ReadLongs/ReadInts decode runs of one-byte values 16 at a time with vectors, without a loop over the run (#24). Mixed data (90% small values) is read 1.48× faster than the plain loop on the i7 and the EPYC, and small values 2.7–3.6× faster; timestamps stay within 3% of the plain loop, as #29's rule requires.
  • Schema parsing uses a pooled JsonDocument and copies out only what a schema keeps (defaults and custom properties) (#103). Parsing a small schema is 2.07× faster than Apache.Avro on the i7 (1.08× before) and allocates 7.2 KB instead of 8.4 KB. Content after the schema JSON is now reported by the JSON parser, with its line and column.
  • The generic reader stores arrays of boolean, int, long, float and double items as the primitives themselves, not one AvroValue each (#23). AsArray() still returns the items as values. The new TryGetInt64Array and the other typed accessors return the memory without a copy, and FromInt64Array and the other typed factories wrap existing memory. The writer takes these arrays from their memory: booleans, floats and doubles as one copy, ints and longs with no kind check per item. On the i7-12800H (.NET 10), reading an array of 1,000 items with the generic model is 4.4× faster for ints (3,241 to 734 ns), 3.5× for longs and 6.2× for doubles, and allocates a quarter to a half as much. A record with 64 long counters reads in 591 ns instead of 748 ns. Against Apache.Avro, array reads are now 6.6–16.6× faster, up from 1.7–2.8×.
  • The array returned by AsArray() for these item types is no longer a List<AvroValue>. Code that cast it to List<AvroValue> has to copy it instead (AsArray() has always been documented as IReadOnlyList<AvroValue>).
  • Generated code is smaller (#114). Nullable unions of a primitive or a record, and arrays and maps of primitives, strings, bytes or records, are one call to an AvroGeneratedCode helper instead of an inlined switch or loop. The union helpers are inlined; the collection helpers take struct codecs (IAvroCodec<T>) that the JIT specializes. The schema-evolution reader is one ReadField switch instead of a case and a helper method per field. For a production-style schema of 12 records, generated code went from 18,653 to 9,193 lines and its assembly from 377 KB to 165 KB; the 31-field test Order record went from 2,200 to 1,038 lines.
  • Generated records with more than 32 fields are serialized by several methods of up to 32 fields each, which keeps each within the JIT's limit on tracked locals (#113). On a 140-field record (WideRecordBenchmarks, i7-12800H, .NET 10), a generated read takes 525 ns, against 661 ns with one method per record and 729 ns before this change; a write takes 374 ns, against 419 and 492 ns.
  • Generated readers no longer allocate each collection twice: they construct records with a constructor that skips the property initializers (#113). Block counts are checked against the input without a 64-bit division, the plan for another writer schema is cached per generated type, int and long arrays are written with the new AvroWriter.WriteInts/WriteLongs, and date converts through DateOnly.DayNumber. A generated read of the benchmark Order record allocates 2,048 B instead of 2,192 B, with times unchanged within noise; reading another version of a schema takes 179 ns (181 to 193 ns before).
  • new T() gives every field with a schema default its default (#112): primitives, strings, bytes, enums, nullable unions whose default is for their first branch, and arrays and maps of those. Before, only reading data that lacked the field applied the default. Code that relied on these fields starting at zero or empty sees the defaults instead.
  • The Apache compatibility mode makes code written for avrogen classes compile unchanged (#116): property names default to the Avro field names, as avrogen's are (AvroSharpPropertyNames=pascal keeps PascalCase), and types have avrogen's static _SCHEMA and an instance Schema (Apache's Avro.Schema). AvroSharp's schema is AvroSharpSchema in that mode, on records as on fixed types. CodeGenOptions.PropertyNaming is now nullable: null means the mode's default.
  • Generated property names title-case segments in capitals (#117): USER_ID becomes UserId (it was USERID) and HTTP2_PORT becomes Http2Port; segments with lower-case letters keep their capitals (txId is TxId). This renames such properties in existing generated code.
  • [GeneratedCode] on generated types carries the generator's package version (for example 0.1.2) instead of 0.0.0.0 (#107): MinVer sets AssemblyVersion to major.0.0.0.
  • Generated code no longer checks a union branch's value for null after a type pattern proved it is not, and null checks have one pair of parentheses (#107).

Fixed

  • Generated code: a line terminator other than \n in a schema's doc text (\r, U+0085, U+2028 or U+2029) ended the generated /// comment, so the rest of the doc was compiled as C#. For a field's doc, that could add members to the generated type. Doc text is now split on every C# line terminator, and other control characters are dropped (#110).
  • Generated readers limited zero-size array items (nulls, empty records) per array, not per value. Arrays nested in arrays multiplied the limit: an 8 KB input could declare about 131 million items and allocate about 2 GB. The limit of 65,536 now covers the whole value, as in the generic reader. Each generated Read starts a new budget, so values read one after another from the same reader, as in a container block, each get the full limit (#109).
  • The generator reports a union whose branches map to the same C# type, for example a uuid string and a uuid fixed (both Guid), as error AVROGEN003. Before, it generated code that didn't compile (CS8120), and whose writer couldn't have chosen the branch anyway. The message suggests AvroSharpLogicalTypes=raw (#108).
  • Generated writers check enum values (#111): a C# enum can hold any number, which was written as an out-of-range ordinal that readers reject. Writing it now throws AvroException naming the field.
  • A bytes decimal with no bytes is rejected with AvroDataException instead of being read as 0, as in Java (#111). The Apache compatibility mode's decimals check the same.
  • Put on a generated record's null-typed field rejects values other than null, which were silently dropped on write (#111).
  • A generated type's Schema is one instance even when first read on several threads at once. Before, a thread that lost the race to parse it could keep its own instance, which missed the cache of resolution plans and the same-schema fast path. The new AvroGeneratedCode.PublishSchema stores the first one.
  • Hostile input (#129):
    • A writer schema that nested arrays, maps or unions between recursive records overflowed the stack, which ends the process: MaxDepth counted only records. Values may now be nested 8 × MaxDepth levels deep (1,024 by default), counting arrays, maps and records, on every path that reads by the writer's schema: the generic and resolving readers, skipped fields, and the transcoder generated types use. The thread's stack is also checked every 16 levels of new depth.
    • A record of many null fields takes no bytes but creates a value per field, so 3 bytes could allocate 160 MB. MaxZeroSizeItems now counts each zero-size item as the values reading it creates: one, plus one per field of each record in it.
    • The bzip2 codec let IndexOutOfRangeException escape on corrupt blocks; bzip2, xz, zstandard and snappy now report every corrupt block as InvalidDataException. A snappy length of 2^31 or more is rejected instead of ending in ArgumentOutOfRangeException.
    • The resolving-reader cache kept every writer schema alive for as long as its reader schema lived, so each file opened with OpenGeneric(stream, readerSchema) leaked its schema.
    • Parsing allocated in proportion to the nesting depth for each field default, because it copied the JSON path; a deeply nested schema allocated 55 times its JSON, and now about 20 times at any depth.
    • A container header may hold at most 1,024 metadata entries, and its avro.schema entry at most AvroFileReaderOptions.MaxSchemaLength bytes (4 MiB by default).
    • Pipelined container reading reserved up to 4 × MaxBlockLength per block; it now reserves at most MaxBlockLength.
    • AvroStreamReader decoded an object again after every read, so a stream that returned one byte per read made a 100 KB string cost 100,000 decodes. A decode cut off inside a string, bytes or fixed value now waits for that value's bytes.
    • createReader of AvroMessageReader and AvroRegistryMessageReader could run more than once per schema when several threads met the schema at once; it runs once, as documented.
  • Schema resolution (#130):
    • A writer union branch was read as the first reader branch of the same unqualified name, although another branch had its full name: com.y.Event in a union with com.x.Event failed to read, or was read as com.x.Event and changed branch when written again. Branches now match by full name (or a reader alias) across the whole union first, then by unqualified name, then by promotion, as in Java.
    • A reader field's alias takes the writer field before a reader field of the writer field's name does, as in Java: the specification defines aliases as rewriting the writer's schema.
    • Container blocks of more than 65,536 zero-size objects (nulls, empty records), which Java and AvroFileWriter write, were rejected. They are now limited by AvroFileReaderOptions.MaxZeroSizeValuesPerBlock (16,777,216 values by default, counted as MaxZeroSizeItems is), and AvroFileWriter starts a new block every 65,536 objects. A block of objects that take at least a byte each can no longer declare more objects than it has bytes.
  • Code generation (#131):
    • new T() gives every field its schema default, as reading data that lacks the field does. Before, defaults with no C# literal were dropped: records, fixed values, logical types (date of 0 became 0001-01-01, uuid became Guid.Empty, a decimal of 0.01 became 0), and collections and unions of them. A fixed default left the field null, so the new value could not be written. Such defaults are now stored as their Avro encoding and decoded by the field's reader.

    • Arrays of enums, fixed values, unions, nullable values, logical types and nested collections are written with a loop over the list on C# 12 and earlier. The loop over the list's span, which needs C# 13 on .NET 9 and 10, made projects pinned to C# 12 fail to compile (CS9202).

    • A float or double default beyond the type's range is float.PositiveInfinity (and the like), not the invalid Infinityf.

    • A type named like a member the generator adds to it (Schema, Read, ToAvroBytes, a fixed type's Size or Value), an enum named like one of its symbols, and a type named var are renamed with a trailing _ and reported as AVROGEN005. Before, they did not compile. In the Apache.Avro compatibility mode, which finds types by name, they are errors (AVROGEN003).

    • Contextual keywords as type or namespace names (record, file, scoped, required, partial, nameof, _ and others) are escaped with @, and generated code no longer uses nameof, which a namespace of that name captured.

    • The source generator failed on two types whose names differ only by case (cs.Order and cs.order) and dropped all its output; they now generate as two types. avrosharp gen rejects them with exit code 1 instead of writing one file over the other on Windows and macOS.

    • AVROGEN003 now also reports:

      • two types that map to the same C# name (-m a:M -m b:M with a.X and b.X);
      • a type whose C# name is also a namespace (app.events next to app.events.Click);
      • an AvroSharpNamespace that is not a C# namespace.

      avrosharp gen --namespace rejects such a namespace as a usage error (exit code 2).

    • Two schema files that need each other's types are reported as a circular reference, naming the other file. Before, both errors only said that a type was not defined.

  • AvroSharp.CodeGen could not be loaded by .NET 8 and 9 applications: its only build, for netstandard2.0, was compiled against the System.Text.Json 10 package, which those runtimes do not have. It now also targets net8.0, which uses the framework's System.Text.Json. The source generator keeps the netstandard2.0 build.

0.1.1 - 2026-09-28

The first complete release of all four packages. 0.1.0's publish stopped partway, so AvroSharp.CodeGen 0.1.0 was never published; use 0.1.1. The code is the same as 0.1.0.

Fixed

  • Releasing: AvroSharp.Generators no longer produces a symbol package. It had no .pdb in it, since the generator ships under analyzers/, and nuget.org's rejection stopped the 0.1.0 publish before AvroSharp.CodeGen. The generator's PDB is embedded in its DLL instead. The release workflow now pushes each package separately and checks, while packing, that every symbol package contains a PDB.

0.1.0 - 2026-09-28

The first preview. It covers:

  • Schemas: parsing, writing, canonical form and fingerprints.
  • The generic data model: binary and JSON encoding, and schema evolution.
  • Code generation from .avsc files.
  • Container files with every codec in the specification.
  • Messages: single-object messages and schema-registry framing.
  • Streams of objects.

Packages: AvroSharp, AvroSharp.Codecs, AvroSharp.CodeGen and AvroSharp.Generators. As a 0.x release, the API may still change before 1.0.

Added

  • Nightly fuzzing (.github/workflows/fuzz.yml): every libFuzzer target runs for 30 minutes a night with SharpFuzz, keeping its corpus between runs and uploading crash inputs. The libFuzzer steps in fuzz/README.md are now verified; a first 20-minute run of all seven targets found no crashes.
  • Generated readers resolve through a plan built once per writer schema (#69): Read(ref reader, writerSchema) reads the writer's fields straight into the type's properties, in the writer's order. Fields of the same schema are read directly, numbers (also arrays of them) are promoted in place, enum ordinals are remapped, writer-only fields are skipped and missing fields take their pre-encoded defaults; only other differences (nested records of another version, unions, logical types) are transcoded, one field at a time. On the evolution benchmark this reads in 387 ns, against 506 ns for the generic resolving reader and 1,033 ns for Apache.Avro (i5-3570K); it was 616 ns.
  • AvroValueTransformer.Transform(schema, value, transform): walks a generic value with its schema, through records, arrays, maps and union branches (resolved as the generic writer resolves them), and calls the transform for each non-null primitive, enum or fixed value inside a record field. The transform gets an AvroFieldContext: the record, the field with its properties (for example confluent:tags), the value's own schema, and the field's FullName as Confluent's data rules name it. Records, arrays and maps are copied only where a value changes, so a transform that changes nothing returns the same instance. A depth limit (128 by default) stops cyclic records. This is the basis for field-level data rules, encryption and redaction (#82).
  • Schema-registry wire framing (AvroSharp.Messages), next to single-object encoding and with no registry client dependency: AvroRegistryFraming for Confluent (0x00 + 4-byte ID, byte-identical to Confluent's serializer; and the version 1 0x01 + GUID framing), Apicurio (4- and 8-byte IDs) and AWS Glue (0x03, compression byte, UUID; zlib-compressed payloads read and written), AvroSchemaId, AvroRegistryMessage for writing, and AvroRegistryMessageReader for reading through an IAvroSchemaIdResolver (a synchronous lookup with an asynchronous fill; AvroSchemaIdStore in memory). Read functions are cached per ID with a last-hit check. ConfluentSchemaIdHeader encodes and decodes Confluent's __key_schema_id/__value_schema_id header values, read with ReadPayload. Hostile input (short or unknown headers, unknown IDs, trailing bytes, corrupt zlib, payloads expanding past MaxPayloadLength) raises AvroDataException; a RegistryMessage fuzz target covers it.
  • Schema references for registries: AvroSchema.ToJson(referencedSchemas) writes the named types of referenced subjects by name, as Java's Schema.toString(referencedSchemas, false) does, for a schema parsed against them (AvroSchemaParser.AddNamedSchemas). Checked against Apache Avro Java 1.12 output and against Confluent Schema Registry 7.7, which stores the text unchanged and finds the registered version when it is registered again.
  • Pipelined container reading: AvroFileReader<T>.ReadAllPipelinedAsync(blocksAhead) reads and decompresses blocks on a background task, up to blocksAhead (2 by default) ahead of the caller, who decodes them. Blocks pass through a bounded System.Threading.Channels channel, so memory stays bounded, and each block's buffers return to the pool once it is decoded. An error in a block is raised when the caller reaches that block, and stopping the enumeration stops the background task. The netstandard targets reference System.Threading.Channels (#32).
  • Streams of objects without a container (AvroSharp.Streams): AvroStreamWriter writes objects one after another in the binary encoding, buffered (AvroStreamOptions.BufferSize, 64 KiB by default), with Write/WriteAsync/Flush/FlushAsync. AvroStreamReader reads them back to the end of the stream with TryRead, ReadAll and ReadAllAsync (IAsyncEnumerable<T>), for generated types or as generic values, optionally resolved to a reader schema. Nothing in the encoding delimits objects, so each is decoded to find its end. An object cut off by the buffer is decoded again once more data is read, and a failure is reported only once the stream has ended or MaxDatumLength (64 MiB by default) bytes are buffered, with the object's stream offset. Objects that encode to no bytes cannot be delimited and are rejected (#32).
  • GenericRecord.TryGetValue(int, out AvroValue): field access by position without an exception for a position the record lacks, alongside the name-based overload.
  • AvroSharp.Codecs: the snappy, zstandard, bzip2 and xz codecs in one package, on fully managed libraries: Snappier (plus System.IO.Hashing for snappy's CRC-32), ZstdSharp.Port, SharpZipLib and Lzma.Net. AvroCodecs.All gives a reader every codec, since a file's codec is not known until it is opened. The writing settings and defaults match Apache Avro Java's: zstandard level 3 with an optional content checksum, bzip2 block size 9, and xz level 6. Snappy blocks end with the big-endian CRC-32 of their data, which is checked on read, and zstandard frames are read with or without a content size or checksum. Damaged blocks, and blocks that decompress past MaxBlockLength, are AvroDataException. Checked against files written by Apache Avro Java 1.12.2 in every codec (tests/TestData/java-avro, with the script that writes them), and Java reads the files these codecs write. Snappy and bzip2 are also checked against Apache.Avro C#'s codec packages in both directions. When a file uses a standard codec that is not available, the reader's error names this package (#32).
  • Seeking and splitting container files, as in the Java implementation: AvroFileReader<T>.PreviousSync (the current block's start), Seek (to a block start), Sync (to the first block after a position, found by scanning for the sync marker) and PastSync (whether the current block belongs to the next split). A split [start, end) is read with Sync(start) and while (TryRead(out var v) && !PastSync(end)). Positions before the header's marker go to the first block, so a marker in the metadata (Apache's syncInMeta.avro) is not mistaken for a block. Needs a seekable stream.
  • Asynchronous container files: AvroFileReader.OpenAsync/OpenGenericAsync and ReadAllAsync (IAsyncEnumerable<T>), and AvroFileWriter<T>.WriteAsync/FlushAsync/DisposeAsync; both types implement IAsyncDisposable. The asynchronous paths do no synchronous I/O (tested with a stream that throws on it); objects are encoded into memory and blocks decoded from memory synchronously. Cancellation is checked before every block. The writer now writes the header with its first block, or when flushed or disposed, instead of when created. Microsoft.Bcl.AsyncInterfaces is referenced for netstandard2.0 (it already came in through System.Text.Json).
  • Single-object encoding (AvroSharp.Messages): AvroMessage writes and reads the C3 01 marker and the writer schema's CRC-64-AVRO fingerprint, and writes whole messages for generic values or any type with a write delegate. AvroMessageReader looks each message's fingerprint up in an IAvroSchemaStore (AvroSchemaStore is a thread-safe in-memory one), creates the read function once per writer schema, and can resolve every message to one reader schema. Headers that are missing, unknown fingerprints and bytes left after the object are AvroDataException. Checked byte for byte against Java's messageV1 test message. Apache.Avro C# has no single-object API, so there is no benchmark baseline.
  • Object container files (AvroSharp.Containers): AvroFileWriter and AvroFileReader for generic values or any type with a write or read delegate (generated types pass their static Write/Read methods). The built-in codecs are null and deflate (BCL, raw DEFLATE); others plug in through AvroCodec. Blocks are written at a configurable sync interval, the header carries application metadata, and a write that throws leaves nothing behind. The reader resolves to an optional reader schema, verifies each block's sync marker, and bounds hostile input: block and metadata sizes are limited (MaxBlockLength, 64 MiB by default, also applied to decompressed data), object counts are checked against the block's bytes, and bytes left after a block's last object are an error. Checked against Apache's weather.avro, weather-sorted.avro (deflate) and syncInMeta.avro, and against Apache.Avro in both directions with random schemas and data.
  • Schema resolution for generated types: Read(ref reader, writerSchema) and FromAvroBytes(data, writerSchema) read data written with another version of the type's schema. Data of the same canonical schema takes the direct path; other data is transcoded, following the resolution rules, straight into the type's own encoding in a reused per-thread buffer and read from there, without creating generic values. On the evolution benchmark this reads 1.65x faster than Apache.Avro's resolving reader with 29% of its allocations (i5-3570K).
  • Schema resolution for the generic model: GenericDatumReader.Create(writerSchema, readerSchema) reads data written with one schema version as another, following the specification. Record fields are matched by name or alias, writer-only fields are skipped (sized blocks in one step), and missing reader fields take their defaults. Named types match by full name, unqualified name or alias. Numbers are promoted, string and bytes convert, and enum symbols are matched with the reader's default for unknown ones. Unions resolve per branch. Incompatible schemas are rejected when the reader is created; a mismatch that only some data would hit (a union branch, an unknown enum symbol without a default) is reported when such a value is read.
  • AvroSchemaParseOptions.AllowIdenticalRedefinitions: a parser may accept a named type that an earlier Parse call defined, when both definitions have the same canonical form; the first definition stays in use. The source generator turns it on, so schema sets that inline shared types in every file (as Apache's one-file-at-a-time tooling requires) generate each type once. Different definitions are still an error that names both files.
  • Generator option AvroSharpPropertyNames=avro (CodeGenOptions.PropertyNaming): keep the Avro field names as property names, as Apache's avrogen does, so code written against avrogen classes compiles unchanged. C# keywords are escaped.
  • Apache.Avro compatibility mode for generated code (AvroSharpApacheCompatible=true, requires a reference to Apache.Avro): records also implement Avro.Specific.ISpecificRecord, fixed types derive from Avro.Specific.SpecificFixed, and logical types use Apache's .NET types, so Apache's SpecificDatumWriter<T>/SpecificDatumReader<T> and AvroSharp's serializers work on the same classes and produce the same bytes. Put also accepts what Apache's reader passes: an enum's ordinal, and an AvroDecimal for a decimal on fixed. The generator reports AVROGEN004 when the property is set without the reference.
  • Logical types in generated code:
    • date becomes DateOnly and time-millis/time-micros become TimeOnly (DateTime/TimeSpan where those types don't exist);
    • timestamp-millis/timestamp-micros become DateTimeOffset, and the local-timestamp variants become DateTime;
    • uuid (on string or fixed(16)) becomes Guid;
    • decimal with a precision up to 28 becomes decimal.
    • Set the MSBuild property AvroSharpLogicalTypes=raw to keep the underlying types.
    • The conversions are public in AvroLogicalValues:
      • decimals are exact, raising an error instead of rounding;
      • times and timestamps are truncated towards negative infinity to the logical type's precision;
      • out-of-range data raises AvroDataException.
  • Generated records implement IAvroSpecificRecord: Schema, Get(int) and Put(int, object?), field access by position following the contract of Apache.Avro's ISpecificRecord without depending on it. Put checks the value's type (no implicit widening, as with Apache's casts) and names the field in errors.
  • Code generation from schema files: the AvroSharp.Generators source generator (an incremental generator for .avsc files passed as AdditionalFiles) and the AvroSharp.CodeGen engine it uses. Records become partial classes with static Write/Read methods (plus ToAvroBytes/FromAvroBytes) that call AvroWriter/AvroReader directly in schema order; enums become C# enums; fixed types become size-checked wrappers. Schema files may refer to each other's named types; errors are reported at the file, line and column. Generated readers enforce the same hostile-input limits as the generic reader. The generator needs the .NET 10 SDK or Visual Studio 2026.
  • JSON encoding for the generic model: GenericDatumJsonWriter and GenericDatumJsonReader, following the specification (wrapped union values, byte strings for bytes and fixed, enum symbols). Record fields may appear in any order and missing fields take their defaults; NaN and infinities are written as strings, as Apache.Avro C# does. Checked both ways against Apache.Avro's JsonEncoder/JsonDecoder.
  • Fuzz targets (fuzz/AvroSharp.Fuzz, SharpFuzz/libFuzzer) for schema parsing and generic binary and JSON data, container files, single-object messages and schema resolution (the transcoder checked against the resolving reader), each also checking round trips. They run on every build as a seeded mutation smoke test.
  • Generic data model (AvroSharp.Generic): AvroValue, a 16-byte struct holding any Avro value without boxing (primitives inline, enums as schema plus ordinal), GenericRecord and GenericFixed. GenericDatumWriter and GenericDatumReader compile a schema once into a cached, thread-safe plan of typed nodes; union branches are selected from the value's kind or schema name. Arrays of int/long/float/double are read in bulk. Hostile input is bounded: block counts are checked against the remaining input (using each record's minimum encoded size), pre-allocation is capped, zero-size items draw from a per-read budget, and record nesting is limited when reading and writing (GenericDatumReaderOptions, GenericDatumWriterOptions).
  • AvroReader.ReadLongs/ReadInts: bulk varint reads; on net8+ a Vector128 check decodes runs of one-byte values 16 at a time.
  • Binary encoding (AvroSharp.IO): AvroWriter writes directly into an IBufferWriter<byte> or a Span<byte>; AvroReader reads from a ReadOnlySpan<byte> or a multi-segment ReadOnlySequence<byte>, returning slices of the input for bytes, string and fixed when contiguous. Bulk double/float array items are a single copy on little-endian hardware. Length prefixes are checked against the remaining input before any allocation; malformed data raises AvroDataException.
  • Schema model (AvroSharp.Schemas): immutable primitive, record, enum, array, map, union and fixed schemas; names, namespaces and aliases; record fields with defaults, order and aliases; custom properties.
  • All Avro 1.12 logical types. Unknown or invalid logical types are ignored and kept as properties, as the specification requires.
  • AvroSchemaParser and AvroSchema.Parse/ParseAsync: System.Text.Json parser over UTF-8 with default-value validation, optional comments, and errors that report the JSON path, line and column. A parser keeps named types across calls, so schemas split over several files can refer to each other.
  • Full schema JSON writer, Parsing Canonical Form, and CRC-64-AVRO, MD5 and SHA-256 fingerprints.
  • Tests against Apache Avro's schema-tests.txt vectors, property-based interop tests against Apache.Avro (C#), a Native AOT smoke test, and schema-parse benchmarks gated against Apache.Avro.
  • AvroCodecNames: the codec names defined by the specification.
  • Repository skeleton: build settings, analyzers, public API tracking, strong naming, TUnit tests on .NET 8/9/10 and .NET Framework 4.8.1, and CI on Linux and Windows (x64 and Arm64), including .NET Framework 4.8.1 on Windows (#72).

Changed

  • AvroSchema.ToJson() now gives the text of Apache Avro Java's Schema.toString(): a named type's aliases come after its custom properties, numbers with a fraction or exponent in defaults and properties are printed as Java prints them (1e10 becomes 1.0E10, computed exactly on every runtime), and only ", `` and control characters are escaped. The parsed schema and its canonical form are unchanged.
  • AvroWriter.WriteString encodes in one pass when the buffer has room for the longest possible encoding. It reserves the prefix for 3 bytes per char, encodes, then writes the real length, moving the bytes down when the prefix is shorter. Otherwise it counts first, as before, so a fixed-size span destination that holds exactly the encoding still works. On the string encode benchmark (i7, ShortRun), it went from 0.88x to 0.83x Apache.Avro's time (#25).
  • Container files use their streams less:
    • The writer writes each block, with its count, size and sync marker, in one stream write instead of three. That is one call or one await per block. Its block buffer is sized from the sync interval (#62).
    • The reader keeps its block buffer between blocks while it is large enough, reuses the buffer that limits decompressed size, and decodes objects from an array segment (#62).
    • PastSync checks the stream's length once per block, not for every object of the documented split loop (#62).
    • A header that needs several fills is still re-scanned, but an attempt that runs out of data no longer allocates. Metadata is recorded as positions and turned into strings and copies once, when the header is complete (#70).
  • The container benchmarks and the Native AOT smoke test cover every codec. The benchmark baselines are Apache.Avro's codec packages (#32).
  • Reading generated types with a writer schema (Read(ref reader, writerSchema), as container files and single-object messages do for every object) no longer compares the two schemas' canonical forms for every record. A writer schema remembers the last schema found to have its canonical form, so repeated checks compare references (#60).
  • Resolving records whose reader field order differs from the writer's no longer allocates an array per record; the slots are rented from the shared array pool (#61).
  • AvroMessageReader checks the last schema used before its fingerprint dictionary, which saves the dictionary lookup when messages repeat one schema (#63).
  • Each package ships its own README (AvroSharp.CodeGen and AvroSharp.Generators no longer show the repository README), and all three include THIRD-PARTY-NOTICES.md (#66).

Fixed

  • THIRD-PARTY-NOTICES.md said the packages contain no third-party code, but they compile in source from Polyfill (MIT). The notice now includes Polyfill's copyright and license, and is packed into every package (#66).
  • Generated code needed C# 9 (new() initializers, ??=, is { } and is not patterns), so it failed to compile in netstandard2.0 and .NET Framework projects, which default to C# 7.3. It now uses constructs every version accepts, and emits nullable annotations only for C# 8 and later.
  • Invalid UTF-8 inside a JSON string (schema JSON or JSON data) raised InvalidOperationException instead of AvroSchemaException/AvroDataException. Found by the fuzz smoke test.