Getting Started

This guide will help you quickly get up and running with Wolfgang.Etl.FixedWidth.

Prerequisites

  • .NET 8.0 SDK or later to follow this guide (the package also targets net462, net481, netstandard2.0, net5.0, net6.0, and net7.0 for older runtimes)

Installation

Via NuGet Package Manager

dotnet add package Wolfgang.Etl.FixedWidth

Via Package Manager Console

Install-Package Wolfgang.Etl.FixedWidth

Quick Start

Define a record class

Map each property to a fixed-width column using [FixedWidthField] attributes:

using Wolfgang.Etl.FixedWidth.Attributes;
using Wolfgang.Etl.FixedWidth.Enums;

public class PersonRecord
{
    [FixedWidthField(0, 10)]
    public string FirstName { get; set; } = string.Empty;



    [FixedWidthField(1, 10)]
    public string LastName { get; set; } = string.Empty;



    [FixedWidthField(2, 3, Alignment = FieldAlignment.Right, Pad = '0')]
    public int Age { get; set; }
}

Extract records from fixed-width text

using System.IO;
using System.Threading;
using Wolfgang.Etl.FixedWidth;

var inputData =
    "Alice     Anderson  025\n" +
    "Bob       Baker     042\n" +
    "Charlie   Clark     033";

var reader = new StringReader(inputData);

var extractor = new FixedWidthExtractor<PersonRecord>(reader);

await foreach (var person in extractor.ExtractAsync(CancellationToken.None))
{
    Console.WriteLine($"{person.FirstName} {person.LastName}, Age {person.Age}");
}

Console.WriteLine($"Total extracted: {extractor.CurrentItemCount}");

Load records to fixed-width text

using System.IO;
using System.Threading;
using Wolfgang.Etl.FixedWidth;

var writer = new StringWriter();

var loader = new FixedWidthLoader<PersonRecord>(writer);

await loader.LoadAsync(recordsAsyncEnumerable, CancellationToken.None);

Console.WriteLine(writer.ToString());
Console.WriteLine($"Total loaded: {loader.CurrentItemCount}");

In production, replace StringReader / StringWriter with a FileStream or StreamReader — the extractor and loader accept any TextReader / TextWriter, or a raw Stream directly.

Controlling line endings

FixedWidthExtractor reads \n, \r, and \r\n automatically, so input needs no configuration.

For output, the loader writes each record with its TextWriter's newline. To force a specific ending regardless of platform — e.g. Unix \n for a downstream mainframe or FTP consumer — pass a TextWriter with the NewLine you want:

using System.IO;

// Force Unix (LF) line endings, even on Windows
await using var stream = File.Create("output.dat");
await using var writer = new StreamWriter(stream) { NewLine = "\n" };
using var loader = new FixedWidthLoader<PersonRecord>(writer);

await loader.LoadAsync(recordsAsyncEnumerable, CancellationToken.None);

NewLine accepts any string; the default is Environment.NewLine.

Checkpoint and resume

For multi-GB files, opt in to byte-offset tracking so a crashed run can resume without re-reading. Persist CurrentByteOffset after each record and pass it back as StartByteOffset on restart:

// First run
await using var stream = File.OpenRead("huge.dat");
using var extractor = new FixedWidthExtractor<Record>(stream) { TrackByteOffset = true };
await foreach (var record in extractor.ExtractAsync(token))
{
    Process(record);
    SaveCheckpoint(extractor.CurrentByteOffset);
}

// Resume after a crash
await using var stream = File.OpenRead("huge.dat");
using var extractor = new FixedWidthExtractor<Record>(stream) { StartByteOffset = LoadCheckpoint() };
await foreach (var record in extractor.ExtractAsync(token)) { /* the remainder */ }

Terminators (\n/\r/\r\n), multi-byte UTF-8, and a leading BOM are counted exactly. Tracking is opt-in and requires the Stream constructor (seekable for resume); the default read path is unchanged. On resume, header lines are not re-skipped and SkipItemCount applies from the resumed position.

Inspecting the layout

FixedWidthSchema.For<T>() exposes the resolved field layout as a read-only view — handy for generating documentation, building validation tooling, or debugging a mapping. It runs the same validation as extraction, so an invalid layout throws here too.

var schema = FixedWidthSchema.For<PersonRecord>();

foreach (var field in schema.Fields)   // includes skip columns (field.IsSkip)
{
    Console.WriteLine($"{field.StartPosition}-{field.EndPosition}  {field.Name}  ({field.Length})");
}

Console.WriteLine($"Line width: {schema.ExpectedLineWidth}, fields: {schema.FieldCount}, skips: {schema.SkipCount}");

Each FixedWidthFieldInfo carries Name, StartPosition/EndPosition, Length, ColumnIndex, PropertyType, Alignment, Pad, Format, Header, and NumberStyles. Skipped columns have IsSkip == true and a SkipMessage.

ToDiagram() renders the layout as a text table for logs, tickets, or docs:

Console.WriteLine(FixedWidthSchema.For<EmployeeRecord>().ToDiagram());
// Position  Field           Type    Length  Align  Pad  Format
// --------  --------------  ------  ------  -----  ---  ------
// 0-9       FirstName       String  10      Left   ' '
// 10-17     [skip]                  8
// 18-23     EmployeeNumber  String  6       Left   ' '
//
// Total width: 24  |  Columns: 3 (2 fields + 1 skip)  |  Delimiter: none

Defining the layout in code

When you can't decorate the record type, or the layout is chosen at runtime, build the schema with FixedWidthSchemaBuilder<T> and assign it to the extractor/loader Schema property instead of using attributes:

var schema = new FixedWidthSchemaBuilder<CustomerRecord>()
    .Field(r => r.CustomerId, index: 0, length: 8)
    .Field(r => r.Name, index: 1, length: 30)
    .Skip(index: 2, length: 5)
    .Field(r => r.Balance, index: 3, length: 9, alignment: FieldAlignment.Right, format: "0000000.00")
    .Build();

using var extractor = new FixedWidthExtractor<CustomerRecord>(reader) { Schema = schema };

The lambda selectors are type-safe (no magic strings) and index is the zero-based column ordinal, matching [FixedWidthField(index, length)]. A built schema is equivalent to an attribute-resolved one — same validation and the same Fields / ToDiagram() introspection — and overrides any attributes on the type when set. See the SchemaBuilder example.

Binary / mainframe records

Mainframe (COBOL) files mix text with binary numeric fields — COMP big-endian integers and COMP-3 packed decimals — and are not newline-delimited: each record is a fixed number of bytes. FixedWidthBinaryExtractor<T> / FixedWidthBinaryLoader<T> read and write them by byte count, so packed/binary bytes that happen to be 0x0A/0x0D are never mistaken for separators. Declare the layout with [FixedWidthBinaryField] (widths in bytes; BinaryFieldType picks the decoding):

public class AccountRecord
{
    [FixedWidthBinaryField(0, 8, BinaryFieldType.Text)]
    public string AccountId { get; set; } = string.Empty;

    [FixedWidthBinaryField(1, 4, BinaryFieldType.Binary)]              // COMP
    public int TransactionCount { get; set; }

    [FixedWidthBinaryField(2, 5, BinaryFieldType.PackedDecimal, Scale = 2)]   // PIC S9(7)V99 COMP-3
    public decimal Balance { get; set; }
}

await using var stream = File.OpenRead("accounts.dat");
using var extractor = new FixedWidthBinaryExtractor<AccountRecord>(stream);
await foreach (var account in extractor.ExtractAsync(CancellationToken.None)) { /* … */ }

Text fields decode with the encoding (ASCII by default). For EBCDIC, register the code-page provider and pass the encoding: Encoding.RegisterProvider(CodePagesEncodingProvider.Instance) then new FixedWidthBinaryExtractor<AccountRecord>(stream, Encoding.GetEncoding("IBM037")). Writing is symmetric via FixedWidthBinaryLoader<T>. See the BinaryRecords example.

Transforming between layouts

FixedWidthTransformer<TSource, TDestination> reformats records from one layout to another (reorder, add/remove, or format-convert fields) as the projection stage between an extractor and a loader:

using var extractor   = new FixedWidthExtractor<LegacyRecord>(sourceReader);
using var transformer = new FixedWidthTransformer<LegacyRecord, ModernRecord>(
    legacy => new ModernRecord { Id = legacy.OldId, Name = legacy.FullName.Trim() });
using var loader      = new FixedWidthLoader<ModernRecord>(destinationWriter);

var modern = transformer.TransformAsync(extractor.ExtractAsync(token), token);
await loader.LoadAsync(modern, token);

When source and destination share property names and compatible types, FixedWidthTransformer<LegacyRecord, ModernRecord>.ByMatchingProperties() builds the copy automatically (the destination needs a public parameterless constructor).

Reading files with multiple record types

Mainframe and EDI batch files interleave several record layouts on different lines — a header, detail rows, and a trailer — distinguished by a discriminator character. FixedWidthMultiRecordExtractor routes each line to the right POCO. Register one rule per type; the first matching predicate wins.

using var extractor = new FixedWidthMultiRecordExtractor(reader)
    .When(line => line[0] == 'H', typeof(HeaderRecord))
    .When(line => line[0] == 'D', typeof(DetailRecord))
    .When(line => line[0] == 'T', typeof(TrailerRecord));

await foreach (var record in extractor.ExtractAsync(token))
{
    switch (record)
    {
        case HeaderRecord h: /* ... */ break;
        case DetailRecord d: /* ... */ break;
        case TrailerRecord t: /* ... */ break;
    }
}

From a file or other Stream, the encoding travels on the options record (the TextReader form above already knows its encoding, so it takes no record):

using var extractor = new FixedWidthMultiRecordExtractor
(
    File.OpenRead("batch.txt"),
    new FixedWidthMultiRecordExtractorOptions { Encoding = Encoding.Latin1 }
)
    .When(line => line[0] == 'H', typeof(HeaderRecord))
    .When(line => line[0] == 'D', typeof(DetailRecord))
    .When(line => line[0] == 'T', typeof(TrailerRecord));

Each record type keeps its own independent [FixedWidthField] layout. A line matching no rule throws by default; set UnmatchedLineHandling = UnmatchedLineHandling.Skip to drop it or register a catch-all with .Otherwise(typeof(UnknownRecord)). The extractor shares the family's HeaderLineCount, FieldDelimiter, ValueParser, SkipItemCount/MaximumItemCount, dead-letter OnError, and progress reporting.

Composing an ETL pipeline

The whole extract → transform → load flow can be written as one fluent chain on the generic EtlPipeline (from Wolfgang.Etl.Abstractions 0.16.0). FixedWidthExtractor<T> source factories hang off EtlPipeline.Create() and FixedWidthLoader<T> sink terminators hang off the pipeline, with the extractor/loader configuration exposed as inline setters:

using Wolfgang.Etl.Abstractions;
using Wolfgang.Etl.FixedWidth;

await EtlPipeline
    .Create()
    .FixedWidthExtractor<PersonRecord>("people.dat")
    .Through(KeepAdults)                 // optional stream-to-stream transform delegate
    .FixedWidthLoader<PersonRecord>("people.txt")
    .WriteHeader(true)
    .FieldDelimiter(" | ")
    .RunAsync();

Every source and sink has path, Stream, and TextReader/TextWriter overloads (plus an existing-FixedWidthExtractor<T> overload). Path factories own the file stream they open and dispose it when the run finishes, on success or failure; caller-supplied streams, readers, and writers are left open. See the PipelineExtensions example for a runnable walk-through.

Metrics and observability

The extractor and loader can emit System.Diagnostics.Metrics instruments from the meter Wolfgang.Etl.FixedWidth — counters (items.extracted, items.loaded, items.skipped, lines.read) and a duration histogram (operation.duration), each tagged with etl.operation and etl.record_type. Metrics are zero-config: they activate automatically when a listener subscribes to the meter — there is no flag to set — so the telemetry flows to Prometheus, Grafana, Application Insights, and so on:

builder.Services.AddOpenTelemetry()
    .WithMetrics(m => m.AddMeter("Wolfgang.Etl.FixedWidth"));

When nothing is listening, the extract/load loop (sampling once per operation) runs no metric code, so there is no overhead. See the Metrics example for a raw MeterListener walk-through.

Next Steps

  • Browse the Examples for more detailed scenarios
  • Explore the API Reference for detailed documentation
  • Read the Introduction to learn more about Wolfgang.Etl.FixedWidth

Common Issues

  • FieldOverflowException during loading — a value exceeds its declared field length. Either increase the Length in [FixedWidthField] or switch to FixedWidthConverter.Truncate.
  • MalformedLineException during extraction — a line is shorter or longer than expected. Set MalformedLineHandling to Skip to ignore such lines, or correct the input data.

Additional Resources