Changelog
All notable changes to this project are documented here. The format follows Keep a Changelog and the project uses Semantic Versioning.
Unreleased
Breaking
The public API review before 1.0 (#134) renames and moves members, so the API can be frozen without them (#73). Regenerate checked-in generated code with the new avrosharp gen: code from 0.2 calls support members that moved. From 1.0, code generated by a version keeps compiling against later runtimes of the same major version (the policy).
- Generated-code support:
AvroGeneratedCode,AvroRecordPlan,AvroPlanCache,AvroConversionandAvroUninitializedmove toAvroSharp.Serialization.Generated.IAvroCodec<T>is nowIAvroValueSerializer<T>, the primitive codecsAvro…Serializer, and the generated nestedAvroCodecstructValueSerializer: "codec" means block compression only.- Removed, because generated code no longer calls them:
GetRecordPlan(writer, reader)without the cache,PutTypeMismatch(object?, string, string),AvroRecordPlan.Target(int)andConversion(int), andIsSameSchema(useAvroSchema.HasSameCanonicalForm).
- Exceptions:
AvroDataException(wasAvroSharp.IO) andAvroSchemaException(wasAvroSharp.Schemas) are inAvroSharp. - Schema lookup:
IAvroSchemaStoreis nowIAvroSchemaResolver, withGetSchemaAsync, which an implementer has to add. TheAvroMessageReaderfactories name its parameterresolver. AvroSerializer:Serialize(output, value)andTrySerialize(destination, value, out bytesWritten)take the output first.- Renamed:
GenericDatumReader.SchemaandGenericDatumJsonReader.Schemaare nowWriterSchema;EnumSchema.Defaultis nowDefaultSymbol;AvroReader.Skip(long)is nowSkipRaw;AvroValue.FromByteArray,FromGenericRecordandFromGenericFixedare nowFromBytes,FromRecordandFromFixed;AvroFileReader<T>.GetMetadataStringis nowTryGetMetadataString(key, out value);CodeGenOptions.DefaultNamespace,NamespaceMappingandPropertyNamingare nowNamespace,NamespaceMapandPropertyNames.
- Changed types:
AvroFileReader<T>.Codecis theAvroCodec(its name isCodec.Name);AvroFileWriterOptions.MetadataholdsReadOnlyMemory<byte>values;SchemaFingerprint.Crc64AvroEmptyis along;ConfluentSchemaIdHeader.Encodetakes anAvroSchemaId.
AvroCodec.CreateDeflateis replaced bynew DeflateCodec(level).- Validation: the generic and schema parse options reject a
MaxDepthbelow 1 and a negativeMaxZeroSizeItems.CodeGenOptionsrejects nullable annotations with aLanguageVersionbelow 8, and any version below 7.
Added
AvroMessageReader<T>.ReadAsync, which fetches an unknown fingerprint through the resolver once, andAvroSchemaStore.GetSchemaAsync(#134).AvroSchema.HasSameCanonicalForm: whether two schemas have the same encoding.Equalsstays reference equality, asAvroSchema's docs now say (#134).DeflateCodec: the built-in codec is public, withDefaultandLevel, as the Codecs package's codecs have (#134).- New overloads:
AvroSerializer.Deserialize<T>(in ReadOnlySequence<byte>, AvroSchema writerSchema);GenericDatumReader.Read(in ReadOnlySequence<byte>);AvroMessage.Write(output, in AvroValue, GenericDatumWriter);AvroRegistryMessage.Write(output, framing, id, in AvroValue, GenericDatumWriter)(#134).
- Default limits as constants:
GenericDatumReaderOptions.DefaultMaxDepthandDefaultMaxZeroSizeItems,GenericDatumWriterOptions.DefaultMaxDepthandAvroSchemaParseOptions.DefaultMaxDepth(#134). - Generator settings (#134):
- The
AvroSharpNamespaceMapMSBuild property,avro.ns:CSharp.Nsentries separated by;, like the CLI's--namespace-map. - Generator properties accept their values in any case.
- A value the generator doesn't recognize is warning AVROGEN006; before, a typo was ignored silently.
--language-version 7implies--no-nullable.
- The
- Records built from another record's fields: a record given a field that belongs to another record takes a copy, so
new RecordSchema(name, other.Fields.Append(field))works. It threw before (#134). - A dev container (
.devcontainer/) andbuild/ci-local.sh, which build and test as the Linux CI job does. The container has Ubuntu 24.04 with the .NET 8, 9 and 10 SDKs, the Native AOT prerequisites, CI's variables and a cached NuGet volume; the script runs the job's steps in order. DocFX and ReportGenerator are pinned local tools (.config/dotnet-tools.json), used by CI, the docs workflow and the container alike (#139).
Fixed
- A fixed type named
EqualsorGetHashCodegenerated code that did not compile (CS0542); it is renamed, with a note, like the other generated member names (#141). - A decimal default that the generated C#
decimalcannot hold (beyond 96 bits, or with more digits than the precision) generated code whose constructor threw; it is now a generation error (#141).
Changed
- Tests for the remaining gaps of the test review (#142): enum and fixed aliases, the resolving reader's options, the JSON writer's widening, depth limit and shape errors, maps of any
IReadOnlyDictionary, hostile container files (overlong varints, truncation on the async path, trailing bytes when pipelined),AvroRegistryMessageReader.ReadAsyncwith a missing schema or a cancelled fetch, logical-value range errors, and control characters in schemas. - The nightly fuzzing workflow also runs a random-schema code-generation test (#141): random schemas, with hostile names and every kind of default, are generated, compiled for C# 7.3, 12 and the latest version, and round-tripped. PR CI runs it on 100 schemas.
- The API reference on the documentation site is built from the net10.0 build, so it shows the .NET 8+ API (
AvroSerializer,IAvroSerializable<T>and the overloads that take no delegates), and lists those members with their equivalents on other targets (#136).
0.2.0 - 2026-09-29
The avrosharp command-line tool; generated types that read older versions of their schema, write and read memory without allocating, and start with their schema defaults; and the fixes from an independent review: bounded memory and nesting for hostile input, schema resolution that follows the specification and Java, and code generation fixes. The source generator needs the .NET 10 SDK or Visual Studio 2026 and later; the generated code and the tool run on .NET 8 and later (and the code on .NET Standard 2.0).
Added
- The
avrosharpcommand-line tool, packageAvroSharp.Tool: adotnet toollike Apache.Avro'savrogen, for .NET 8 and later (#33).avrosharp genwrites the source generator's code for.avscfiles, or folders of them, with the generator's options: one file per type, in folders for its namespace, or with--flatin one folder.avrosharp schema canonicalandschema fingerprintprint the Parsing Canonical Form, and the CRC-64-AVRO, MD5 or SHA-256 fingerprint as hex, base64 or (CRC-64) the decimal Java prints.- Files may refer to each other's types in any order. Errors are in the compiler's format, with the generator's diagnostic IDs. The exit code is 0 on success, 1 when the command fails and 2 for an invalid command line.
CodeGenOptions.NamespaceMapping(--namespace-mapin the tool): C# namespaces for Avro namespaces and the namespaces under them, as avrogen's--namespace a:b. Only the generated C# namespace changes, not the schema.GeneratedSource.Namespacegives each file's C# namespace.SchemaFileSetinAvroSharp.CodeGen: parses schema files that refer to each other's named types, in any order, as the source generator and the tool do.- Bulk booleans (#26):
AvroReader.ReadBooleanschecks and copies a block of booleans as one span, andAvroWriter.WriteBooleanswrites one as a single copy. Generated code uses them forbooleanarrays, and the generic reader and writer use them for boolean arrays. - Generated records implement new interfaces (#115):
IAvroWritable(WriteTo(ref AvroWriter)) andIAvroReadable(ReadFrom(ref AvroReader), which fills an existing instance and reuses its lists, dictionaries and records), and on .NET 8 and laterIAvroSerializable<T>with the staticSchema,WriteandReadmembers. With it,AvroSerializer.Serialize/TrySerialize/Deserialize,AvroFileWriter.Create<T>(stream),AvroFileReader.Open<T>(stream)/OpenAsync<T>,AvroStreamWriter.Create<T>,AvroStreamReader.Open<T>,AvroMessage.ToArray<T>/Write<T>,AvroMessageReader.Create<T>,AvroRegistryMessage.ToArray<T>/Write<T>andAvroRegistryMessageReader.Create<T>need no delegates. Readers of files and messages resolve other versions of the type's schema. - Generated records write and read memory without allocating (#115):
TryWriteAvroBytes(Span<byte>, out int)(no exception when the value does not fit),WriteAvroBytes(Span<byte>),WriteAvroBytes(IBufferWriter<byte>),FromAvroBytes(data, out int bytesConsumed)andFromAvroBytes(in ReadOnlySequence<byte>). AvroWriter.WriteIntsandWriteLongswrite array items in bulk.CodeGenOptions.LanguageVersion: the C# version generated code may use. The source generator sets it from the project; with C# 11 or later, generated code addsIAvroSerializable<T>for .NET 8 and later.- Generated code is easier to debug and read (#117): records have a
[DebuggerDisplay]with their first three simple fields, the serializers are[DebuggerNonUserCode], and a property whose logical type keeps its raw type (a decimal wider than 28 digits,duration,timestamp-nanos, or any logical type withAvroSharpLogicalTypes=raw) says in its documentation what the value means. - Fixed types have
==,!=andAsSpan()(#117). - The generator reports a property renamed to avoid a clash (for example
user_idanduserIdin one record) as informational diagnostic AVROGEN005, naming both fields;GeneratedSource.Notescarries these notes for other callers ofCSharpCodeGenerator(#117). - Generated files suppress CS8981, so all-lower-case Avro type names compile without warnings (#117).
AvroFileReaderOptions.MaxSchemaLengthandMaxZeroSizeValuesPerBlock(#129, #130).GenericDatumJsonReader.ReadDefault: converts a schema default to anAvroValue, the value a reader gives a field that the data lacks (#131).
Changed
- Packages list only the dependencies they use (#20).
AvroSharp.GeneratorsandAvroSharp.CodeGendepend only on AvroSharp, andAvroSharp.Codecsonly on AvroSharp and its compression libraries; before, they also listed System.Memory, System.Text.Json and other packages that AvroSharp brings. AvroSharp's netstandard2.1 build no longer depends on System.Memory or Microsoft.Bcl.AsyncInterfaces, which that target has built in. The generator still needs the .NET 10 SDK or Visual Studio 2026 and later;docs/design.mdrecords what fails on older SDKs. - Faster varint encode (#102). Where PDEP is fast (Intel from Haswell, AMD from Zen 3), 3- to 8-byte values are one inline 8-byte store; elsewhere one out-of-line call writes 3 to 5 bytes with one 4-byte store. On the i7-12800H, 3- to 8-byte encodes are 28–38% faster and 3-byte encode is 1.52× faster than Apache.Avro (1.09× before); on the EPYC 7543 they are 25–35% faster, and on an i5-3570K and a Ryzen 5 3500U without fast PDEP, 3 to 5 bytes are 7–33% faster. 1-byte encode is 6–8% slower on the fast-PDEP machines. Record writes are unchanged or faster; generated writes are 7–8% faster on the i7 and the EPYC.
AvroWriter.WriteLongs/WriteIntskeep the write position in a local while the buffer has room, instead of storing and reloading it after every value (#27, the per-value round trip through memory that limited one-at-a-time writes on the EPYC). They are at least as fast as single writes for every length on the machines measured, and 1-byte bulk writes are 19–48% faster than single ones.- Bulk
ReadLongs/ReadIntsdecode runs of one-byte values 16 at a time with vectors, without a loop over the run (#24). Mixed data (90% small values) is read 1.48× faster than the plain loop on the i7 and the EPYC, and small values 2.7–3.6× faster; timestamps stay within 3% of the plain loop, as #29's rule requires. - Schema parsing uses a pooled
JsonDocumentand copies out only what a schema keeps (defaults and custom properties) (#103). Parsing a small schema is 2.07× faster than Apache.Avro on the i7 (1.08× before) and allocates 7.2 KB instead of 8.4 KB. Content after the schema JSON is now reported by the JSON parser, with its line and column. - The generic reader stores arrays of
boolean,int,long,floatanddoubleitems as the primitives themselves, not oneAvroValueeach (#23).AsArray()still returns the items as values. The newTryGetInt64Arrayand the other typed accessors return the memory without a copy, andFromInt64Arrayand the other typed factories wrap existing memory. The writer takes these arrays from their memory: booleans, floats and doubles as one copy, ints and longs with no kind check per item. On the i7-12800H (.NET 10), reading an array of 1,000 items with the generic model is 4.4× faster for ints (3,241 to 734 ns), 3.5× for longs and 6.2× for doubles, and allocates a quarter to a half as much. A record with 64 long counters reads in 591 ns instead of 748 ns. Against Apache.Avro, array reads are now 6.6–16.6× faster, up from 1.7–2.8×. - The array returned by
AsArray()for these item types is no longer aList<AvroValue>. Code that cast it toList<AvroValue>has to copy it instead (AsArray()has always been documented asIReadOnlyList<AvroValue>). - Generated code is smaller (#114). Nullable unions of a primitive or a record, and arrays and maps of primitives, strings, bytes or records, are one call to an
AvroGeneratedCodehelper instead of an inlined switch or loop. The union helpers are inlined; the collection helpers take struct codecs (IAvroCodec<T>) that the JIT specializes. The schema-evolution reader is oneReadFieldswitch instead of a case and a helper method per field. For a production-style schema of 12 records, generated code went from 18,653 to 9,193 lines and its assembly from 377 KB to 165 KB; the 31-field testOrderrecord went from 2,200 to 1,038 lines. - Generated records with more than 32 fields are serialized by several methods of up to 32 fields each, which keeps each within the JIT's limit on tracked locals (#113). On a 140-field record (
WideRecordBenchmarks, i7-12800H, .NET 10), a generated read takes 525 ns, against 661 ns with one method per record and 729 ns before this change; a write takes 374 ns, against 419 and 492 ns. - Generated readers no longer allocate each collection twice: they construct records with a constructor that skips the property initializers (#113). Block counts are checked against the input without a 64-bit division, the plan for another writer schema is cached per generated type,
intandlongarrays are written with the newAvroWriter.WriteInts/WriteLongs, anddateconverts throughDateOnly.DayNumber. A generated read of the benchmarkOrderrecord allocates 2,048 B instead of 2,192 B, with times unchanged within noise; reading another version of a schema takes 179 ns (181 to 193 ns before). new T()gives every field with a schema default its default (#112): primitives, strings, bytes, enums, nullable unions whose default is for their first branch, and arrays and maps of those. Before, only reading data that lacked the field applied the default. Code that relied on these fields starting at zero or empty sees the defaults instead.- The Apache compatibility mode makes code written for
avrogenclasses compile unchanged (#116): property names default to the Avro field names, as avrogen's are (AvroSharpPropertyNames=pascalkeeps PascalCase), and types have avrogen's static_SCHEMAand an instanceSchema(Apache'sAvro.Schema). AvroSharp's schema isAvroSharpSchemain that mode, on records as on fixed types.CodeGenOptions.PropertyNamingis now nullable:nullmeans the mode's default. - Generated property names title-case segments in capitals (#117):
USER_IDbecomesUserId(it wasUSERID) andHTTP2_PORTbecomesHttp2Port; segments with lower-case letters keep their capitals (txIdisTxId). This renames such properties in existing generated code. [GeneratedCode]on generated types carries the generator's package version (for example0.1.2) instead of0.0.0.0(#107): MinVer setsAssemblyVersiontomajor.0.0.0.- Generated code no longer checks a union branch's value for null after a type pattern proved it is not, and null checks have one pair of parentheses (#107).
Fixed
- Generated code: a line terminator other than
\nin a schema'sdoctext (\r, U+0085, U+2028 or U+2029) ended the generated///comment, so the rest of the doc was compiled as C#. For a field's doc, that could add members to the generated type. Doc text is now split on every C# line terminator, and other control characters are dropped (#110). - Generated readers limited zero-size array items (
nulls, empty records) per array, not per value. Arrays nested in arrays multiplied the limit: an 8 KB input could declare about 131 million items and allocate about 2 GB. The limit of 65,536 now covers the whole value, as in the generic reader. Each generatedReadstarts a new budget, so values read one after another from the same reader, as in a container block, each get the full limit (#109). - The generator reports a union whose branches map to the same C# type, for example a
uuidstring and auuidfixed (bothGuid), as error AVROGEN003. Before, it generated code that didn't compile (CS8120), and whose writer couldn't have chosen the branch anyway. The message suggestsAvroSharpLogicalTypes=raw(#108). - Generated writers check enum values (#111): a C# enum can hold any number, which was written as an out-of-range ordinal that readers reject. Writing it now throws
AvroExceptionnaming the field. - A
bytesdecimal with no bytes is rejected withAvroDataExceptioninstead of being read as 0, as in Java (#111). The Apache compatibility mode's decimals check the same. Puton a generated record'snull-typed field rejects values other thannull, which were silently dropped on write (#111).- A generated type's
Schemais one instance even when first read on several threads at once. Before, a thread that lost the race to parse it could keep its own instance, which missed the cache of resolution plans and the same-schema fast path. The newAvroGeneratedCode.PublishSchemastores the first one. - Hostile input (#129):
- A writer schema that nested arrays, maps or unions between recursive records overflowed the stack, which ends the process:
MaxDepthcounted only records. Values may now be nested 8 ×MaxDepthlevels deep (1,024 by default), counting arrays, maps and records, on every path that reads by the writer's schema: the generic and resolving readers, skipped fields, and the transcoder generated types use. The thread's stack is also checked every 16 levels of new depth. - A record of many
nullfields takes no bytes but creates a value per field, so 3 bytes could allocate 160 MB.MaxZeroSizeItemsnow counts each zero-size item as the values reading it creates: one, plus one per field of each record in it. - The bzip2 codec let
IndexOutOfRangeExceptionescape on corrupt blocks; bzip2, xz, zstandard and snappy now report every corrupt block asInvalidDataException. A snappy length of 2^31 or more is rejected instead of ending inArgumentOutOfRangeException. - The resolving-reader cache kept every writer schema alive for as long as its reader schema lived, so each file opened with
OpenGeneric(stream, readerSchema)leaked its schema. - Parsing allocated in proportion to the nesting depth for each field default, because it copied the JSON path; a deeply nested schema allocated 55 times its JSON, and now about 20 times at any depth.
- A container header may hold at most 1,024 metadata entries, and its
avro.schemaentry at mostAvroFileReaderOptions.MaxSchemaLengthbytes (4 MiB by default). - Pipelined container reading reserved up to 4 ×
MaxBlockLengthper block; it now reserves at mostMaxBlockLength. AvroStreamReaderdecoded an object again after every read, so a stream that returned one byte per read made a 100 KB string cost 100,000 decodes. A decode cut off inside a string, bytes or fixed value now waits for that value's bytes.createReaderofAvroMessageReaderandAvroRegistryMessageReadercould run more than once per schema when several threads met the schema at once; it runs once, as documented.
- A writer schema that nested arrays, maps or unions between recursive records overflowed the stack, which ends the process:
- Schema resolution (#130):
- A writer union branch was read as the first reader branch of the same unqualified name, although another branch had its full name:
com.y.Eventin a union withcom.x.Eventfailed to read, or was read ascom.x.Eventand changed branch when written again. Branches now match by full name (or a reader alias) across the whole union first, then by unqualified name, then by promotion, as in Java. - A reader field's alias takes the writer field before a reader field of the writer field's name does, as in Java: the specification defines aliases as rewriting the writer's schema.
- Container blocks of more than 65,536 zero-size objects (
nulls, empty records), which Java andAvroFileWriterwrite, were rejected. They are now limited byAvroFileReaderOptions.MaxZeroSizeValuesPerBlock(16,777,216 values by default, counted asMaxZeroSizeItemsis), andAvroFileWriterstarts a new block every 65,536 objects. A block of objects that take at least a byte each can no longer declare more objects than it has bytes.
- A writer union branch was read as the first reader branch of the same unqualified name, although another branch had its full name:
- Code generation (#131):
new T()gives every field its schema default, as reading data that lacks the field does. Before, defaults with no C# literal were dropped: records, fixed values, logical types (dateof 0 became 0001-01-01,uuidbecameGuid.Empty, a decimal of 0.01 became 0), and collections and unions of them. A fixed default left the fieldnull, so the new value could not be written. Such defaults are now stored as their Avro encoding and decoded by the field's reader.Arrays of enums, fixed values, unions, nullable values, logical types and nested collections are written with a loop over the list on C# 12 and earlier. The loop over the list's span, which needs C# 13 on .NET 9 and 10, made projects pinned to C# 12 fail to compile (CS9202).
A float or double default beyond the type's range is
float.PositiveInfinity(and the like), not the invalidInfinityf.A type named like a member the generator adds to it (
Schema,Read,ToAvroBytes, a fixed type'sSizeorValue), an enum named like one of its symbols, and a type namedvarare renamed with a trailing_and reported as AVROGEN005. Before, they did not compile. In the Apache.Avro compatibility mode, which finds types by name, they are errors (AVROGEN003).Contextual keywords as type or namespace names (
record,file,scoped,required,partial,nameof,_and others) are escaped with@, and generated code no longer usesnameof, which a namespace of that name captured.The source generator failed on two types whose names differ only by case (
cs.Orderandcs.order) and dropped all its output; they now generate as two types.avrosharp genrejects them with exit code 1 instead of writing one file over the other on Windows and macOS.AVROGEN003 now also reports:
- two types that map to the same C# name (
-m a:M -m b:Mwitha.Xandb.X); - a type whose C# name is also a namespace (
app.eventsnext toapp.events.Click); - an
AvroSharpNamespacethat is not a C# namespace.
avrosharp gen --namespacerejects such a namespace as a usage error (exit code 2).- two types that map to the same C# name (
Two schema files that need each other's types are reported as a circular reference, naming the other file. Before, both errors only said that a type was not defined.
AvroSharp.CodeGencould not be loaded by .NET 8 and 9 applications: its only build, for netstandard2.0, was compiled against the System.Text.Json 10 package, which those runtimes do not have. It now also targets net8.0, which uses the framework's System.Text.Json. The source generator keeps the netstandard2.0 build.
0.1.1 - 2026-09-28
The first complete release of all four packages. 0.1.0's publish stopped partway, so AvroSharp.CodeGen 0.1.0 was never published; use 0.1.1. The code is the same as 0.1.0.
Fixed
- Releasing:
AvroSharp.Generatorsno longer produces a symbol package. It had no.pdbin it, since the generator ships underanalyzers/, and nuget.org's rejection stopped the 0.1.0 publish beforeAvroSharp.CodeGen. The generator's PDB is embedded in its DLL instead. The release workflow now pushes each package separately and checks, while packing, that every symbol package contains a PDB.
0.1.0 - 2026-09-28
The first preview. It covers:
- Schemas: parsing, writing, canonical form and fingerprints.
- The generic data model: binary and JSON encoding, and schema evolution.
- Code generation from
.avscfiles. - Container files with every codec in the specification.
- Messages: single-object messages and schema-registry framing.
- Streams of objects.
Packages: AvroSharp, AvroSharp.Codecs, AvroSharp.CodeGen and AvroSharp.Generators. As a 0.x release, the API may still change before 1.0.
Added
- Nightly fuzzing (
.github/workflows/fuzz.yml): every libFuzzer target runs for 30 minutes a night with SharpFuzz, keeping its corpus between runs and uploading crash inputs. The libFuzzer steps infuzz/README.mdare now verified; a first 20-minute run of all seven targets found no crashes. - Generated readers resolve through a plan built once per writer schema (#69):
Read(ref reader, writerSchema)reads the writer's fields straight into the type's properties, in the writer's order. Fields of the same schema are read directly, numbers (also arrays of them) are promoted in place, enum ordinals are remapped, writer-only fields are skipped and missing fields take their pre-encoded defaults; only other differences (nested records of another version, unions, logical types) are transcoded, one field at a time. On the evolution benchmark this reads in 387 ns, against 506 ns for the generic resolving reader and 1,033 ns for Apache.Avro (i5-3570K); it was 616 ns. AvroValueTransformer.Transform(schema, value, transform): walks a generic value with its schema, through records, arrays, maps and union branches (resolved as the generic writer resolves them), and calls the transform for each non-null primitive, enum or fixed value inside a record field. The transform gets anAvroFieldContext: the record, the field with its properties (for exampleconfluent:tags), the value's own schema, and the field'sFullNameas Confluent's data rules name it. Records, arrays and maps are copied only where a value changes, so a transform that changes nothing returns the same instance. A depth limit (128 by default) stops cyclic records. This is the basis for field-level data rules, encryption and redaction (#82).- Schema-registry wire framing (
AvroSharp.Messages), next to single-object encoding and with no registry client dependency:AvroRegistryFramingfor Confluent (0x00+ 4-byte ID, byte-identical to Confluent's serializer; and the version 10x01+ GUID framing), Apicurio (4- and 8-byte IDs) and AWS Glue (0x03, compression byte, UUID; zlib-compressed payloads read and written),AvroSchemaId,AvroRegistryMessagefor writing, andAvroRegistryMessageReaderfor reading through anIAvroSchemaIdResolver(a synchronous lookup with an asynchronous fill;AvroSchemaIdStorein memory). Read functions are cached per ID with a last-hit check.ConfluentSchemaIdHeaderencodes and decodes Confluent's__key_schema_id/__value_schema_idheader values, read withReadPayload. Hostile input (short or unknown headers, unknown IDs, trailing bytes, corrupt zlib, payloads expanding pastMaxPayloadLength) raisesAvroDataException; aRegistryMessagefuzz target covers it. - Schema references for registries:
AvroSchema.ToJson(referencedSchemas)writes the named types of referenced subjects by name, as Java'sSchema.toString(referencedSchemas, false)does, for a schema parsed against them (AvroSchemaParser.AddNamedSchemas). Checked against Apache Avro Java 1.12 output and against Confluent Schema Registry 7.7, which stores the text unchanged and finds the registered version when it is registered again. - Pipelined container reading:
AvroFileReader<T>.ReadAllPipelinedAsync(blocksAhead)reads and decompresses blocks on a background task, up toblocksAhead(2 by default) ahead of the caller, who decodes them. Blocks pass through a boundedSystem.Threading.Channelschannel, so memory stays bounded, and each block's buffers return to the pool once it is decoded. An error in a block is raised when the caller reaches that block, and stopping the enumeration stops the background task. The netstandard targets referenceSystem.Threading.Channels(#32). - Streams of objects without a container (
AvroSharp.Streams):AvroStreamWriterwrites objects one after another in the binary encoding, buffered (AvroStreamOptions.BufferSize, 64 KiB by default), withWrite/WriteAsync/Flush/FlushAsync.AvroStreamReaderreads them back to the end of the stream withTryRead,ReadAllandReadAllAsync(IAsyncEnumerable<T>), for generated types or as generic values, optionally resolved to a reader schema. Nothing in the encoding delimits objects, so each is decoded to find its end. An object cut off by the buffer is decoded again once more data is read, and a failure is reported only once the stream has ended orMaxDatumLength(64 MiB by default) bytes are buffered, with the object's stream offset. Objects that encode to no bytes cannot be delimited and are rejected (#32). GenericRecord.TryGetValue(int, out AvroValue): field access by position without an exception for a position the record lacks, alongside the name-based overload.AvroSharp.Codecs: the snappy, zstandard, bzip2 and xz codecs in one package, on fully managed libraries: Snappier (plusSystem.IO.Hashingfor snappy's CRC-32), ZstdSharp.Port, SharpZipLib and Lzma.Net.AvroCodecs.Allgives a reader every codec, since a file's codec is not known until it is opened. The writing settings and defaults match Apache Avro Java's: zstandard level 3 with an optional content checksum, bzip2 block size 9, and xz level 6. Snappy blocks end with the big-endian CRC-32 of their data, which is checked on read, and zstandard frames are read with or without a content size or checksum. Damaged blocks, and blocks that decompress pastMaxBlockLength, areAvroDataException. Checked against files written by Apache Avro Java 1.12.2 in every codec (tests/TestData/java-avro, with the script that writes them), and Java reads the files these codecs write. Snappy and bzip2 are also checked against Apache.Avro C#'s codec packages in both directions. When a file uses a standard codec that is not available, the reader's error names this package (#32).- Seeking and splitting container files, as in the Java implementation:
AvroFileReader<T>.PreviousSync(the current block's start),Seek(to a block start),Sync(to the first block after a position, found by scanning for the sync marker) andPastSync(whether the current block belongs to the next split). A split[start, end)is read withSync(start)andwhile (TryRead(out var v) && !PastSync(end)). Positions before the header's marker go to the first block, so a marker in the metadata (Apache'ssyncInMeta.avro) is not mistaken for a block. Needs a seekable stream. - Asynchronous container files:
AvroFileReader.OpenAsync/OpenGenericAsyncandReadAllAsync(IAsyncEnumerable<T>), andAvroFileWriter<T>.WriteAsync/FlushAsync/DisposeAsync; both types implementIAsyncDisposable. The asynchronous paths do no synchronous I/O (tested with a stream that throws on it); objects are encoded into memory and blocks decoded from memory synchronously. Cancellation is checked before every block. The writer now writes the header with its first block, or when flushed or disposed, instead of when created.Microsoft.Bcl.AsyncInterfacesis referenced for netstandard2.0 (it already came in through System.Text.Json). - Single-object encoding (
AvroSharp.Messages):AvroMessagewrites and reads theC3 01marker and the writer schema's CRC-64-AVRO fingerprint, and writes whole messages for generic values or any type with a write delegate.AvroMessageReaderlooks each message's fingerprint up in anIAvroSchemaStore(AvroSchemaStoreis a thread-safe in-memory one), creates the read function once per writer schema, and can resolve every message to one reader schema. Headers that are missing, unknown fingerprints and bytes left after the object areAvroDataException. Checked byte for byte against Java'smessageV1test message. Apache.Avro C# has no single-object API, so there is no benchmark baseline. - Object container files (
AvroSharp.Containers):AvroFileWriterandAvroFileReaderfor generic values or any type with a write or read delegate (generated types pass their staticWrite/Readmethods). The built-in codecs arenullanddeflate(BCL, raw DEFLATE); others plug in throughAvroCodec. Blocks are written at a configurable sync interval, the header carries application metadata, and a write that throws leaves nothing behind. The reader resolves to an optional reader schema, verifies each block's sync marker, and bounds hostile input: block and metadata sizes are limited (MaxBlockLength, 64 MiB by default, also applied to decompressed data), object counts are checked against the block's bytes, and bytes left after a block's last object are an error. Checked against Apache'sweather.avro,weather-sorted.avro(deflate) andsyncInMeta.avro, and against Apache.Avro in both directions with random schemas and data. - Schema resolution for generated types:
Read(ref reader, writerSchema)andFromAvroBytes(data, writerSchema)read data written with another version of the type's schema. Data of the same canonical schema takes the direct path; other data is transcoded, following the resolution rules, straight into the type's own encoding in a reused per-thread buffer and read from there, without creating generic values. On the evolution benchmark this reads 1.65x faster than Apache.Avro's resolving reader with 29% of its allocations (i5-3570K). - Schema resolution for the generic model:
GenericDatumReader.Create(writerSchema, readerSchema)reads data written with one schema version as another, following the specification. Record fields are matched by name or alias, writer-only fields are skipped (sized blocks in one step), and missing reader fields take their defaults. Named types match by full name, unqualified name or alias. Numbers are promoted,stringandbytesconvert, and enum symbols are matched with the reader's default for unknown ones. Unions resolve per branch. Incompatible schemas are rejected when the reader is created; a mismatch that only some data would hit (a union branch, an unknown enum symbol without a default) is reported when such a value is read. AvroSchemaParseOptions.AllowIdenticalRedefinitions: a parser may accept a named type that an earlierParsecall defined, when both definitions have the same canonical form; the first definition stays in use. The source generator turns it on, so schema sets that inline shared types in every file (as Apache's one-file-at-a-time tooling requires) generate each type once. Different definitions are still an error that names both files.- Generator option
AvroSharpPropertyNames=avro(CodeGenOptions.PropertyNaming): keep the Avro field names as property names, as Apache'savrogendoes, so code written against avrogen classes compiles unchanged. C# keywords are escaped. - Apache.Avro compatibility mode for generated code (
AvroSharpApacheCompatible=true, requires a reference to Apache.Avro): records also implementAvro.Specific.ISpecificRecord, fixed types derive fromAvro.Specific.SpecificFixed, and logical types use Apache's .NET types, so Apache'sSpecificDatumWriter<T>/SpecificDatumReader<T>and AvroSharp's serializers work on the same classes and produce the same bytes.Putalso accepts what Apache's reader passes: an enum's ordinal, and anAvroDecimalfor a decimal on fixed. The generator reportsAVROGEN004when the property is set without the reference. - Logical types in generated code:
datebecomesDateOnlyandtime-millis/time-microsbecomeTimeOnly(DateTime/TimeSpanwhere those types don't exist);timestamp-millis/timestamp-microsbecomeDateTimeOffset, and thelocal-timestampvariants becomeDateTime;uuid(onstringorfixed(16)) becomesGuid;decimalwith a precision up to 28 becomesdecimal.- Set the MSBuild property
AvroSharpLogicalTypes=rawto keep the underlying types. - The conversions are public in
AvroLogicalValues:- decimals are exact, raising an error instead of rounding;
- times and timestamps are truncated towards negative infinity to the logical type's precision;
- out-of-range data raises
AvroDataException.
- Generated records implement
IAvroSpecificRecord:Schema,Get(int)andPut(int, object?), field access by position following the contract of Apache.Avro'sISpecificRecordwithout depending on it.Putchecks the value's type (no implicit widening, as with Apache's casts) and names the field in errors. - Code generation from schema files: the
AvroSharp.Generatorssource generator (an incremental generator for.avscfiles passed asAdditionalFiles) and theAvroSharp.CodeGenengine it uses. Records become partial classes with staticWrite/Readmethods (plusToAvroBytes/FromAvroBytes) that callAvroWriter/AvroReaderdirectly in schema order; enums become C# enums; fixed types become size-checked wrappers. Schema files may refer to each other's named types; errors are reported at the file, line and column. Generated readers enforce the same hostile-input limits as the generic reader. The generator needs the .NET 10 SDK or Visual Studio 2026. - JSON encoding for the generic model:
GenericDatumJsonWriterandGenericDatumJsonReader, following the specification (wrapped union values, byte strings forbytesandfixed, enum symbols). Record fields may appear in any order and missing fields take their defaults; NaN and infinities are written as strings, as Apache.Avro C# does. Checked both ways against Apache.Avro'sJsonEncoder/JsonDecoder. - Fuzz targets (
fuzz/AvroSharp.Fuzz, SharpFuzz/libFuzzer) for schema parsing and generic binary and JSON data, container files, single-object messages and schema resolution (the transcoder checked against the resolving reader), each also checking round trips. They run on every build as a seeded mutation smoke test. - Generic data model (
AvroSharp.Generic):AvroValue, a 16-byte struct holding any Avro value without boxing (primitives inline, enums as schema plus ordinal),GenericRecordandGenericFixed.GenericDatumWriterandGenericDatumReadercompile a schema once into a cached, thread-safe plan of typed nodes; union branches are selected from the value's kind or schema name. Arrays ofint/long/float/doubleare read in bulk. Hostile input is bounded: block counts are checked against the remaining input (using each record's minimum encoded size), pre-allocation is capped, zero-size items draw from a per-read budget, and record nesting is limited when reading and writing (GenericDatumReaderOptions,GenericDatumWriterOptions). AvroReader.ReadLongs/ReadInts: bulk varint reads; on net8+ aVector128check decodes runs of one-byte values 16 at a time.- Binary encoding (
AvroSharp.IO):AvroWriterwrites directly into anIBufferWriter<byte>or aSpan<byte>;AvroReaderreads from aReadOnlySpan<byte>or a multi-segmentReadOnlySequence<byte>, returning slices of the input forbytes,stringandfixedwhen contiguous. Bulkdouble/floatarray items are a single copy on little-endian hardware. Length prefixes are checked against the remaining input before any allocation; malformed data raisesAvroDataException. - Schema model (
AvroSharp.Schemas): immutable primitive, record, enum, array, map, union and fixed schemas; names, namespaces and aliases; record fields with defaults, order and aliases; custom properties. - All Avro 1.12 logical types. Unknown or invalid logical types are ignored and kept as properties, as the specification requires.
AvroSchemaParserandAvroSchema.Parse/ParseAsync:System.Text.Jsonparser over UTF-8 with default-value validation, optional comments, and errors that report the JSON path, line and column. A parser keeps named types across calls, so schemas split over several files can refer to each other.- Full schema JSON writer, Parsing Canonical Form, and CRC-64-AVRO, MD5 and SHA-256 fingerprints.
- Tests against Apache Avro's
schema-tests.txtvectors, property-based interop tests against Apache.Avro (C#), a Native AOT smoke test, and schema-parse benchmarks gated against Apache.Avro. AvroCodecNames: the codec names defined by the specification.- Repository skeleton: build settings, analyzers, public API tracking, strong naming, TUnit tests on .NET 8/9/10 and .NET Framework 4.8.1, and CI on Linux and Windows (x64 and Arm64), including .NET Framework 4.8.1 on Windows (#72).
Changed
AvroSchema.ToJson()now gives the text of Apache Avro Java'sSchema.toString(): a named type'saliasescome after its custom properties, numbers with a fraction or exponent in defaults and properties are printed as Java prints them (1e10becomes1.0E10, computed exactly on every runtime), and only", `` and control characters are escaped. The parsed schema and its canonical form are unchanged.AvroWriter.WriteStringencodes in one pass when the buffer has room for the longest possible encoding. It reserves the prefix for 3 bytes per char, encodes, then writes the real length, moving the bytes down when the prefix is shorter. Otherwise it counts first, as before, so a fixed-size span destination that holds exactly the encoding still works. On the string encode benchmark (i7, ShortRun), it went from 0.88x to 0.83x Apache.Avro's time (#25).- Container files use their streams less:
- The writer writes each block, with its count, size and sync marker, in one stream write instead of three. That is one call or one await per block. Its block buffer is sized from the sync interval (#62).
- The reader keeps its block buffer between blocks while it is large enough, reuses the buffer that limits decompressed size, and decodes objects from an array segment (#62).
PastSyncchecks the stream's length once per block, not for every object of the documented split loop (#62).- A header that needs several fills is still re-scanned, but an attempt that runs out of data no longer allocates. Metadata is recorded as positions and turned into strings and copies once, when the header is complete (#70).
- The container benchmarks and the Native AOT smoke test cover every codec. The benchmark baselines are Apache.Avro's codec packages (#32).
- Reading generated types with a writer schema (
Read(ref reader, writerSchema), as container files and single-object messages do for every object) no longer compares the two schemas' canonical forms for every record. A writer schema remembers the last schema found to have its canonical form, so repeated checks compare references (#60). - Resolving records whose reader field order differs from the writer's no longer allocates an array per record; the slots are rented from the shared array pool (#61).
AvroMessageReaderchecks the last schema used before its fingerprint dictionary, which saves the dictionary lookup when messages repeat one schema (#63).- Each package ships its own README (
AvroSharp.CodeGenandAvroSharp.Generatorsno longer show the repository README), and all three includeTHIRD-PARTY-NOTICES.md(#66).
Fixed
THIRD-PARTY-NOTICES.mdsaid the packages contain no third-party code, but they compile in source from Polyfill (MIT). The notice now includes Polyfill's copyright and license, and is packed into every package (#66).- Generated code needed C# 9 (
new()initializers,??=,is { }andis notpatterns), so it failed to compile in netstandard2.0 and .NET Framework projects, which default to C# 7.3. It now uses constructs every version accepts, and emits nullable annotations only for C# 8 and later. - Invalid UTF-8 inside a JSON string (schema JSON or JSON data) raised
InvalidOperationExceptioninstead ofAvroSchemaException/AvroDataException. Found by the fuzz smoke test.