Getting Started
This guide will help you quickly get up and running with Wolfgang.Etl.FixedWidth.
Prerequisites
- .NET 8.0 SDK or later to follow this guide (the package also targets net462, net481, netstandard2.0, net5.0, net6.0, and net7.0 for older runtimes)
Installation
Via NuGet Package Manager
dotnet add package Wolfgang.Etl.FixedWidth
Via Package Manager Console
Install-Package Wolfgang.Etl.FixedWidth
Quick Start
Define a record class
Map each property to a fixed-width column using [FixedWidthField] attributes:
using Wolfgang.Etl.FixedWidth.Attributes;
using Wolfgang.Etl.FixedWidth.Enums;
public class PersonRecord
{
[FixedWidthField(0, 10)]
public string FirstName { get; set; } = string.Empty;
[FixedWidthField(1, 10)]
public string LastName { get; set; } = string.Empty;
[FixedWidthField(2, 3, Alignment = FieldAlignment.Right, Pad = '0')]
public int Age { get; set; }
}
Extract records from fixed-width text
using System.IO;
using System.Threading;
using Wolfgang.Etl.FixedWidth;
var inputData =
"Alice Anderson 025\n" +
"Bob Baker 042\n" +
"Charlie Clark 033";
var reader = new StringReader(inputData);
var extractor = new FixedWidthExtractor<PersonRecord>(reader);
await foreach (var person in extractor.ExtractAsync(CancellationToken.None))
{
Console.WriteLine($"{person.FirstName} {person.LastName}, Age {person.Age}");
}
Console.WriteLine($"Total extracted: {extractor.CurrentItemCount}");
Load records to fixed-width text
using System.IO;
using System.Threading;
using Wolfgang.Etl.FixedWidth;
var writer = new StringWriter();
var loader = new FixedWidthLoader<PersonRecord>(writer);
await loader.LoadAsync(recordsAsyncEnumerable, CancellationToken.None);
Console.WriteLine(writer.ToString());
Console.WriteLine($"Total loaded: {loader.CurrentItemCount}");
In production, replace StringReader / StringWriter with a FileStream or StreamReader — the extractor and loader accept any TextReader / TextWriter, or a raw Stream directly.
Controlling line endings
FixedWidthExtractor reads \n, \r, and \r\n automatically, so input needs no configuration.
For output, the loader writes each record with its TextWriter's newline. To force a specific ending regardless of platform — e.g. Unix \n for a downstream mainframe or FTP consumer — pass a TextWriter with the NewLine you want:
using System.IO;
// Force Unix (LF) line endings, even on Windows
await using var stream = File.Create("output.dat");
await using var writer = new StreamWriter(stream) { NewLine = "\n" };
using var loader = new FixedWidthLoader<PersonRecord>(writer);
await loader.LoadAsync(recordsAsyncEnumerable, CancellationToken.None);
NewLine accepts any string; the default is Environment.NewLine.
Checkpoint and resume
For multi-GB files, opt in to byte-offset tracking so a crashed run can resume without re-reading. Persist CurrentByteOffset after each record and pass it back as StartByteOffset on restart:
// First run
await using var stream = File.OpenRead("huge.dat");
using var extractor = new FixedWidthExtractor<Record>(stream) { TrackByteOffset = true };
await foreach (var record in extractor.ExtractAsync(token))
{
Process(record);
SaveCheckpoint(extractor.CurrentByteOffset);
}
// Resume after a crash
await using var stream = File.OpenRead("huge.dat");
using var extractor = new FixedWidthExtractor<Record>(stream) { StartByteOffset = LoadCheckpoint() };
await foreach (var record in extractor.ExtractAsync(token)) { /* the remainder */ }
Terminators (\n/\r/\r\n), multi-byte UTF-8, and a leading BOM are counted exactly. Tracking is opt-in and requires the Stream constructor (seekable for resume); the default read path is unchanged. On resume, header lines are not re-skipped and SkipItemCount applies from the resumed position.
Inspecting the layout
FixedWidthSchema.For<T>() exposes the resolved field layout as a read-only view — handy for generating documentation, building validation tooling, or debugging a mapping. It runs the same validation as extraction, so an invalid layout throws here too.
var schema = FixedWidthSchema.For<PersonRecord>();
foreach (var field in schema.Fields) // includes skip columns (field.IsSkip)
{
Console.WriteLine($"{field.StartPosition}-{field.EndPosition} {field.Name} ({field.Length})");
}
Console.WriteLine($"Line width: {schema.ExpectedLineWidth}, fields: {schema.FieldCount}, skips: {schema.SkipCount}");
Each FixedWidthFieldInfo carries Name, StartPosition/EndPosition, Length, ColumnIndex, PropertyType, Alignment, Pad, Format, Header, and NumberStyles. Skipped columns have IsSkip == true and a SkipMessage.
ToDiagram() renders the layout as a text table for logs, tickets, or docs:
Console.WriteLine(FixedWidthSchema.For<EmployeeRecord>().ToDiagram());
// Position Field Type Length Align Pad Format
// -------- -------------- ------ ------ ----- --- ------
// 0-9 FirstName String 10 Left ' '
// 10-17 [skip] 8
// 18-23 EmployeeNumber String 6 Left ' '
//
// Total width: 24 | Columns: 3 (2 fields + 1 skip) | Delimiter: none
Defining the layout in code
When you can't decorate the record type, or the layout is chosen at runtime, build the schema with FixedWidthSchemaBuilder<T> and assign it to the extractor/loader Schema property instead of using attributes:
var schema = new FixedWidthSchemaBuilder<CustomerRecord>()
.Field(r => r.CustomerId, index: 0, length: 8)
.Field(r => r.Name, index: 1, length: 30)
.Skip(index: 2, length: 5)
.Field(r => r.Balance, index: 3, length: 9, alignment: FieldAlignment.Right, format: "0000000.00")
.Build();
using var extractor = new FixedWidthExtractor<CustomerRecord>(reader) { Schema = schema };
The lambda selectors are type-safe (no magic strings) and index is the zero-based column ordinal, matching [FixedWidthField(index, length)]. A built schema is equivalent to an attribute-resolved one — same validation and the same Fields / ToDiagram() introspection — and overrides any attributes on the type when set. See the SchemaBuilder example.
Binary / mainframe records
Mainframe (COBOL) files mix text with binary numeric fields — COMP big-endian integers and COMP-3 packed decimals — and are not newline-delimited: each record is a fixed number of bytes. FixedWidthBinaryExtractor<T> / FixedWidthBinaryLoader<T> read and write them by byte count, so packed/binary bytes that happen to be 0x0A/0x0D are never mistaken for separators. Declare the layout with [FixedWidthBinaryField] (widths in bytes; BinaryFieldType picks the decoding):
public class AccountRecord
{
[FixedWidthBinaryField(0, 8, BinaryFieldType.Text)]
public string AccountId { get; set; } = string.Empty;
[FixedWidthBinaryField(1, 4, BinaryFieldType.Binary)] // COMP
public int TransactionCount { get; set; }
[FixedWidthBinaryField(2, 5, BinaryFieldType.PackedDecimal, Scale = 2)] // PIC S9(7)V99 COMP-3
public decimal Balance { get; set; }
}
await using var stream = File.OpenRead("accounts.dat");
using var extractor = new FixedWidthBinaryExtractor<AccountRecord>(stream);
await foreach (var account in extractor.ExtractAsync(CancellationToken.None)) { /* … */ }
Text fields decode with the encoding (ASCII by default). For EBCDIC, register the code-page provider and pass the encoding: Encoding.RegisterProvider(CodePagesEncodingProvider.Instance) then new FixedWidthBinaryExtractor<AccountRecord>(stream, Encoding.GetEncoding("IBM037")). Writing is symmetric via FixedWidthBinaryLoader<T>. See the BinaryRecords example.
Transforming between layouts
FixedWidthTransformer<TSource, TDestination> reformats records from one layout to another (reorder, add/remove, or format-convert fields) as the projection stage between an extractor and a loader:
using var extractor = new FixedWidthExtractor<LegacyRecord>(sourceReader);
using var transformer = new FixedWidthTransformer<LegacyRecord, ModernRecord>(
legacy => new ModernRecord { Id = legacy.OldId, Name = legacy.FullName.Trim() });
using var loader = new FixedWidthLoader<ModernRecord>(destinationWriter);
var modern = transformer.TransformAsync(extractor.ExtractAsync(token), token);
await loader.LoadAsync(modern, token);
When source and destination share property names and compatible types, FixedWidthTransformer<LegacyRecord, ModernRecord>.ByMatchingProperties() builds the copy automatically (the destination needs a public parameterless constructor).
Reading files with multiple record types
Mainframe and EDI batch files interleave several record layouts on different lines — a header, detail rows, and a trailer — distinguished by a discriminator character. FixedWidthMultiRecordExtractor routes each line to the right POCO. Register one rule per type; the first matching predicate wins.
using var extractor = new FixedWidthMultiRecordExtractor(reader)
.When(line => line[0] == 'H', typeof(HeaderRecord))
.When(line => line[0] == 'D', typeof(DetailRecord))
.When(line => line[0] == 'T', typeof(TrailerRecord));
await foreach (var record in extractor.ExtractAsync(token))
{
switch (record)
{
case HeaderRecord h: /* ... */ break;
case DetailRecord d: /* ... */ break;
case TrailerRecord t: /* ... */ break;
}
}
From a file or other Stream, the encoding travels on the options record (the TextReader form above already knows its encoding, so it takes no record):
using var extractor = new FixedWidthMultiRecordExtractor
(
File.OpenRead("batch.txt"),
new FixedWidthMultiRecordExtractorOptions { Encoding = Encoding.Latin1 }
)
.When(line => line[0] == 'H', typeof(HeaderRecord))
.When(line => line[0] == 'D', typeof(DetailRecord))
.When(line => line[0] == 'T', typeof(TrailerRecord));
Each record type keeps its own independent [FixedWidthField] layout. A line matching no rule throws by default; set UnmatchedLineHandling = UnmatchedLineHandling.Skip to drop it or register a catch-all with .Otherwise(typeof(UnknownRecord)). The extractor shares the family's HeaderLineCount, FieldDelimiter, ValueParser, SkipItemCount/MaximumItemCount, dead-letter OnError, and progress reporting.
Composing an ETL pipeline
The whole extract → transform → load flow can be written as one fluent chain on the generic EtlPipeline (from Wolfgang.Etl.Abstractions 0.16.0). FixedWidthExtractor<T> source factories hang off EtlPipeline.Create() and FixedWidthLoader<T> sink terminators hang off the pipeline, with the extractor/loader configuration exposed as inline setters:
using Wolfgang.Etl.Abstractions;
using Wolfgang.Etl.FixedWidth;
await EtlPipeline
.Create()
.FixedWidthExtractor<PersonRecord>("people.dat")
.Through(KeepAdults) // optional stream-to-stream transform delegate
.FixedWidthLoader<PersonRecord>("people.txt")
.WriteHeader(true)
.FieldDelimiter(" | ")
.RunAsync();
Every source and sink has path, Stream, and TextReader/TextWriter overloads (plus an existing-FixedWidthExtractor<T> overload). Path factories own the file stream they open and dispose it when the run finishes, on success or failure; caller-supplied streams, readers, and writers are left open. See the PipelineExtensions example for a runnable walk-through.
Metrics and observability
The extractor and loader can emit System.Diagnostics.Metrics instruments from the meter Wolfgang.Etl.FixedWidth — counters (items.extracted, items.loaded, items.skipped, lines.read) and a duration histogram (operation.duration), each tagged with etl.operation and etl.record_type. Metrics are zero-config: they activate automatically when a listener subscribes to the meter — there is no flag to set — so the telemetry flows to Prometheus, Grafana, Application Insights, and so on:
builder.Services.AddOpenTelemetry()
.WithMetrics(m => m.AddMeter("Wolfgang.Etl.FixedWidth"));
When nothing is listening, the extract/load loop (sampling once per operation) runs no metric code, so there is no overhead. See the Metrics example for a raw MeterListener walk-through.
Next Steps
- Browse the Examples for more detailed scenarios
- Explore the API Reference for detailed documentation
- Read the Introduction to learn more about Wolfgang.Etl.FixedWidth
Common Issues
- FieldOverflowException during loading — a value exceeds its declared field length. Either increase the
Lengthin[FixedWidthField]or switch toFixedWidthConverter.Truncate. - MalformedLineException during extraction — a line is shorter or longer than expected. Set
MalformedLineHandlingtoSkipto ignore such lines, or correct the input data.