Class FixedWidthExtractor<TRecord>

Namespace
Wolfgang.Etl.FixedWidth
Assembly
Wolfgang.Etl.FixedWidth.dll

Reads a fixed-width text file and yields records of type TRecord as an asynchronous stream.

public class FixedWidthExtractor<TRecord> : ExtractorBase<TRecord, FixedWidthReport>, IExtractWithProgressAndCancellationAsync<TRecord, FixedWidthReport>, IExtractWithCancellationAsync<TRecord>, IExtractWithProgressAsync<TRecord, FixedWidthReport>, IExtractAsync<TRecord>, IReportsItemErrors, IAsyncDisposable, IDisposable where TRecord : notnull, new()

Type Parameters

TRecord

The POCO type representing a single record. Properties decorated with FixedWidthFieldAttribute are populated from each line. The type must have a public parameterless constructor.

Inheritance
ExtractorBase<TRecord, FixedWidthReport>
FixedWidthExtractor<TRecord>
Implements
IExtractWithProgressAndCancellationAsync<TRecord, FixedWidthReport>
IExtractWithCancellationAsync<TRecord>
IExtractWithProgressAsync<TRecord, FixedWidthReport>
IExtractAsync<TRecord>
IReportsItemErrors
Inherited Members
ExtractorBase<TRecord, FixedWidthReport>.ExtractAsync()
ExtractorBase<TRecord, FixedWidthReport>.CreateProgressReport()
ExtractorBase<TRecord, FixedWidthReport>.IncrementCurrentItemCount()
ExtractorBase<TRecord, FixedWidthReport>.IncrementCurrentSkippedItemCount()
ExtractorBase<TRecord, FixedWidthReport>.OnItemError(ItemErrorContext)
ExtractorBase<TRecord, FixedWidthReport>.HandleItemError(ItemErrorContext)
ExtractorBase<TRecord, FixedWidthReport>.DisposeAsync()
ExtractorBase<TRecord, FixedWidthReport>.Dispose()
ExtractorBase<TRecord, FixedWidthReport>.StartedAt
ExtractorBase<TRecord, FixedWidthReport>.Elapsed
ExtractorBase<TRecord, FixedWidthReport>.ReportingInterval
ExtractorBase<TRecord, FixedWidthReport>.CurrentItemCount
ExtractorBase<TRecord, FixedWidthReport>.CurrentSkippedItemCount
ExtractorBase<TRecord, FixedWidthReport>.CurrentErrorItemCount
ExtractorBase<TRecord, FixedWidthReport>.MaximumItemCount
ExtractorBase<TRecord, FixedWidthReport>.SkipItemCount
ExtractorBase<TRecord, FixedWidthReport>.WorkerResilience
ExtractorBase<TRecord, FixedWidthReport>.ErrorPolicy

Examples

// Stream-based (preferred for files — 64 KB buffer reduces syscall overhead):
await using var stream = File.OpenRead("data.txt");
using var extractor = new FixedWidthExtractor<CustomerRecord>(stream);

// TextReader-based (caller owns the reader):
var extractor = new FixedWidthExtractor<CustomerRecord>(reader);

Remarks

Two construction modes are supported, each with different ownership semantics:

  • TextReader constructor — the caller owns the TextReader lifetime. The extractor does not dispose it. Calling Dispose() is optional and has no effect.
  • Stream constructor — the extractor creates an internal StreamReader with a 64 KB buffer for improved throughput on large files. The caller retains ownership of the Stream (it is not closed), but Dispose() must be called to release the internal reader.

Constructors

FixedWidthExtractor(Stream, ILogger<FixedWidthExtractor<TRecord>>, Encoding)

Initializes a new FixedWidthExtractor<TRecord> from a Stream decoded with the supplied Encoding, with diagnostic logging.

[Obsolete("Use the constructor that takes FixedWidthExtractorStreamOptions. This overload will be removed in a future release.")]
public FixedWidthExtractor(Stream stream, ILogger<FixedWidthExtractor<TRecord>> logger, Encoding encoding)

Parameters

stream Stream

The stream to use.

logger ILogger<FixedWidthExtractor<TRecord>>

The logger to use for diagnostic output.

encoding Encoding

The encoding to decode with, or null for the documented default.

Exceptions

ArgumentNullException

stream is null.

FixedWidthExtractor(Stream, Encoding)

Initializes a new FixedWidthExtractor<TRecord> from a Stream decoded with the supplied Encoding.

[Obsolete("Use the constructor that takes FixedWidthExtractorStreamOptions. This overload will be removed in a future release.")]
public FixedWidthExtractor(Stream stream, Encoding encoding)

Parameters

stream Stream

The stream to use.

encoding Encoding

The encoding to decode with, or null for the documented default.

Exceptions

ArgumentNullException

stream is null.

FixedWidthExtractor(Stream, FixedWidthExtractorStreamOptions<TRecord>?, ILogger<FixedWidthExtractor<TRecord>>?)

Initializes a new FixedWidthExtractor<TRecord> over the specified Stream, with the logger as the trailing optional parameter.

public FixedWidthExtractor(Stream stream, FixedWidthExtractorStreamOptions<TRecord>? options = null, ILogger<FixedWidthExtractor<TRecord>>? logger = null)

Parameters

stream Stream

The Stream to use.

options FixedWidthExtractorStreamOptions<TRecord>

Options that control behaviour, including the Encoding to use. When null, the documented defaults apply.

logger ILogger<FixedWidthExtractor<TRecord>>

An optional logger for diagnostic output. When null — or omitted — Instance is used and logging is disabled.

Exceptions

ArgumentNullException

stream is null.

FixedWidthExtractor(TextReader, ILogger<FixedWidthExtractor<TRecord>>?)

Initializes a new FixedWidthExtractor<TRecord> that reads from the specified TextReader with diagnostic logging.

public FixedWidthExtractor(TextReader reader, ILogger<FixedWidthExtractor<TRecord>>? logger = null)

Parameters

reader TextReader

The TextReader to read fixed-width records from.

logger ILogger<FixedWidthExtractor<TRecord>>

The logger instance for diagnostic output.

Exceptions

ArgumentNullException

reader or logger is null.

FixedWidthExtractor(TextReader, FixedWidthExtractorOptions<TRecord>?, ILogger<FixedWidthExtractor<TRecord>>?)

Initializes a new instance that reads from reader with the given configuration.

[SuppressMessage("ApiDesign", "RS0026:Do not add multiple public overloads with optional parameters", Justification = "Shipped shape: the Stream overload keeps its optional options/logger for source compatibility until the constructor set is settled in the 2026-12-15 removal wave (#373 / #343).")]
public FixedWidthExtractor(TextReader reader, FixedWidthExtractorOptions<TRecord>? options = null, ILogger<FixedWidthExtractor<TRecord>>? logger = null)

Parameters

reader TextReader

The reader supplying the fixed-width lines. The caller owns it.

options FixedWidthExtractorOptions<TRecord>

The parsing configuration. null (the default) keeps every default. The reader has already decoded its bytes, so this is the base record without an Encoding.

logger ILogger<FixedWidthExtractor<TRecord>>

An optional logger.

Exceptions

ArgumentNullException

reader is null.

Properties

BlankLineHandling

Specifies what happens when a truly blank line (zero length) is encountered in the file. Evaluated before the skip budget and Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount.

public BlankLineHandling BlankLineHandling { get; set; }

Property Value

BlankLineHandling

Remarks

  • ThrowException (default) — always throws a LineTooShortException regardless of position.
  • Skip — the line is invisible to all counting logic. Does not count toward Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.SkipItemCount or Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount.
  • ReturnDefault — a default TRecord instance is yielded. Counts toward the skip budget if within Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.SkipItemCount, otherwise counts toward Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount.

Note: a line consisting entirely of spaces is not blank — it is a valid data line that will parse to a record with all whitespace-trimmed (empty/default) fields.

LineFilter is not invoked for blank lines.

CurrentByteOffset

The byte offset of the start of the next unread line — the value to persist as a checkpoint after processing each record. Advances as every physical line is consumed (including headers, blank, and skipped lines) so a resume continues exactly where reading stopped.

public long CurrentByteOffset { get; }

Property Value

long

Remarks

Thread-safe: read with Interlocked so it may be sampled from a progress timer thread.

Exceptions

InvalidOperationException

Byte-offset tracking is not enabled — set TrackByteOffset or StartByteOffset first, and construct the extractor from a Stream.

CurrentFilteredLineCount

The number of physical lines read that did not produce a record and were neither skipped by the budget nor rejected: header lines, the separator line, blank lines dropped per BlankLineHandling, lines dropped by LineFilter, and the line that triggered early termination. With CurrentLineNumber this closes the line accounting: CurrentLineNumber = CurrentItemCount + CurrentSkippedItemCount + CurrentRejectedItemCount + CurrentFilteredLineCount.

public int CurrentFilteredLineCount { get; }

Property Value

int

CurrentLineNumber

The 1-based physical line number of the line most recently read from the file. Updated before each line is parsed so that if an exception is thrown, this value points to the offending line. Matches the line number shown in a text editor — no adjustment is needed for header or separator lines.

public long CurrentLineNumber { get; }

Property Value

long

Remarks

Thread-safe: reads are performed with Interlocked so this property may be sampled from a progress-reporting timer thread without a data race.

CurrentRejectedItemCount

The number of parsed records rejected so far: records discarded via Skip and records rejected by RecordValidator. Distinct from Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.CurrentSkippedItemCount (the SkipItemCount pagination budget).

public int CurrentRejectedItemCount { get; }

Property Value

int

FieldDelimiter

The delimiter string present between fields in the source file, or null (default) for pure fixed-width input with no delimiter. Must match the FieldDelimiter used when the file was written.

public string? FieldDelimiter { get; set; }

Property Value

string

Examples

// File was written with FieldDelimiter = " | " — set the same value on the extractor:
// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
FieldDelimiter = " | ",

// Pure fixed-width file with no delimiter (default):
FieldDelimiter = null,

Remarks

When set, the extractor accounts for the delimiter width when calculating field start positions, ensuring each field is read from the correct offset.

FieldSeparator

When non-null, the line immediately following the last header line is treated as a separator and skipped. Has no effect if HeaderLineCount is 0. The value of the character is not used for parsing — only its presence matters. Set to null (default) for no separator. Mirrors FieldSeparator.

public char? FieldSeparator { get; set; }

Property Value

char?

Examples

// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
HeaderLineCount      = 1,
FieldSeparator = '-',  // skips a "----------" separator line after the header
FieldSeparator = null, // no separator line (default)

HasHeader

Convenience property — when set to true, sets HeaderLineCount to 1. When set to false, sets HeaderLineCount to 0. Returns true if HeaderLineCount is greater than zero. Mirrors WriteHeader.

public bool HasHeader { get; set; }

Property Value

bool

Examples

// Skip one header line before reading records:
// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
HeaderLineCount = 1,

// Skip one header line and one separator line:
HeaderLineCount    = 1,
FieldSeparator = '-',

HeaderLineCount

The number of header lines to skip at the beginning of the file before extracting records. Defaults to 0. For the common single-header case, use HasHeader instead.

public int HeaderLineCount { get; set; }

Property Value

int

Examples

// File has two header lines followed by data:
// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
HeaderLineCount = 2,

// Equivalent shorthand for the common single-header case:
HeaderLineCount = 1,

Remarks

Lines 1 through HeaderLineCount are skipped entirely without parsing. If FieldSeparator is also set, the line immediately after the last header line is additionally skipped as a separator line.

LineFilter

A delegate invoked for every data line (after header and separator lines have been skipped) before any parsing occurs. Return Process to parse the line normally, Skip to skip it, or Stop to end the stream immediately without parsing the line.

public Func<string, LineAction> LineFilter { get; set; }

Property Value

Func<string, LineAction>

Examples

// Footer string — stop when a known marker line is reached
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => line == "END" ? LineAction.Stop : LineAction.Process }

// Trailing separator — stop when a line consists entirely of dashes
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => line.All(c => c == '-') ? LineAction.Stop : LineAction.Process }

// EOF marker — stop when a line starts with a sentinel prefix
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => line.StartsWith("$$") ? LineAction.Stop : LineAction.Process }

// Comment lines — skip lines that begin with '#'
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => line.StartsWith("#") ? LineAction.Skip : LineAction.Process }

// Blank line as terminator — stop at the first empty line
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => string.IsNullOrWhiteSpace(line) ? LineAction.Stop : LineAction.Process }

Remarks

Evaluated after BlankLineHandling — blank lines never reach the filter. Evaluated before the skip budget and Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount. Both Skip and Stop are invisible to all counting logic — they do not affect Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.SkipItemCount, Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount, or CurrentSkippedItemCount. Defaults to a function that always returns Process.

MalformedLineHandling

Specifies what happens when a line is encountered that is too short or whose field values cannot be converted to the target property type.

public MalformedLineHandling MalformedLineHandling { get; set; }

Property Value

MalformedLineHandling

Remarks

Defaults to ThrowException. When set to Skip, the line is skipped and CurrentSkippedItemCount is incremented. When set to ReturnDefault, a default instance of TRecord is yielded for the offending line.

OnError

An optional dead-letter sink invoked once for each record that fails to parse (#29). The failed line is reported as a FixedWidthError — its 1-based source line number, raw content, and exception — so the caller can log it, push it to a dead-letter queue, or collect it for post-run inspection. Set MalformedLineHandling to Skip to capture-and-continue; with the default ThrowException the failure is still reported here before the exception is re-thrown. Business rejects from RecordValidator are not reported here — they are not parse errors.

public Action<FixedWidthError>? OnError { get; set; }

Property Value

Action<FixedWidthError>

Examples

var errors = new List<FixedWidthError>();
var extractor = new FixedWidthExtractor<Record>(reader, new FixedWidthExtractorOptions<Record>
{
    MalformedLineHandling = MalformedLineHandling.Skip,
    OnError = errors.Add,
});
await foreach (var ok in extractor.ExtractAsync(token)) { /* only good records */ }
// errors now holds the dead letters

RecordValidator

An optional callback invoked for each fully parsed record, after ValueParser and field assignment but before the record is yielded. Return Accept() to yield it, Skip(string?) to drop it (increments CurrentSkippedItemCount), or Stop(string?) to end extraction. null (the default) applies no validation.

public Func<TRecord, ValidationResult>? RecordValidator { get; set; }

Property Value

Func<TRecord, ValidationResult>

Examples

// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
RecordValidator = record =>
    record.Balance < 0
        ? ValidationResult.Skip("Negative balance")
        : ValidationResult.Accept(),

Schema

An optional layout that overrides the [FixedWidthField] / [FixedWidthSkip] attributes on TRecord (#23). Build one with FixedWidthSchemaBuilder<T> to map a type you cannot decorate, or to define the layout in code. When null (the default) the attribute-based layout is used. The schema's RecordType must be TRecord.

public FixedWidthSchema? Schema { get; set; }

Property Value

FixedWidthSchema

StartByteOffset

The byte offset to seek to before extraction begins — a checkpoint saved from a prior run's CurrentByteOffset. Defaults to 0 (start of stream). A non-zero value requires a seekable Stream-constructed extractor and implicitly enables TrackByteOffset. On resume, header lines are not re-skipped (the checkpoint is past them); Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.SkipItemCount is applied from the resumed position.

public long StartByteOffset { get; set; }

Property Value

long

Exceptions

ArgumentOutOfRangeException

The value is negative.

TrackByteOffset

Enables byte-offset tracking for checkpoint/resume (#31). When true, CurrentByteOffset reports the byte position of the next unread line so a pipeline can save a checkpoint after each record and resume via StartByteOffset after a crash. Defaults to false — tracking is opt-in because it wraps the reader in a byte-counting decoder, changing the read path. Requires the Stream constructor; a caller-owned TextReader has no addressable byte stream. Setting StartByteOffset to a non-zero value enables tracking implicitly.

public bool TrackByteOffset { get; set; }

Property Value

bool

Remarks

Byte offsets are computed with the encoding passed to the Stream constructor (UTF8 by default), so tracking assumes the stream is actually encoded that way. A byte-order mark that would switch the internal StreamReader to a different encoding (for example a UTF-16 BOM when UTF-8 was specified) is not supported for tracking and would yield incorrect offsets — pass the stream's real encoding to the constructor when enabling tracking. A matching-encoding BOM (e.g. a UTF-8 BOM with the default encoding) is handled correctly.

ValueParser

A delegate that converts a raw string read from the file into the target property type. The FieldContext provides the property type, format string, and other field metadata needed to perform the conversion. Defaults to DefaultParser. Mirrors ValueConverter.

public FixedWidthValueParser ValueParser { get; set; }

Property Value

FixedWidthValueParser

Examples

// Treat "Y"/"N" as bool, fall back to DefaultParser for everything else:
new FixedWidthExtractorOptions<TRecord>
{
    ValueParser = (text, ctx) =>
        ctx.PropertyType == typeof(bool)
            ? (object)(text.Span.SequenceEqual("Y".AsSpan()))
            : FixedWidthConverter.DefaultParser(text, ctx),
}

// Parse a custom date format for a specific field:
new FixedWidthExtractorOptions<TRecord>
{
    ValueParser = (text, ctx) =>
        ctx.PropertyName == "BirthDate"
            ? DateTime.ParseExact(text.ToString(), "dd/MM/yyyy", CultureInfo.InvariantCulture)
            : FixedWidthConverter.DefaultParser(text, ctx),
}

Remarks

The delegate must return a value that is assignable to the property's CLR type. Returning an incompatible type will cause an InvalidCastException when the framework attempts to set the property value.

Exceptions

FieldConversionException

The default DefaultParser wraps any parse failure in a FieldConversionException. Custom parsers should do the same so that MalformedLineHandling can handle them uniformly.

Methods

CreateProgressReport()

Creates a progress report snapshot for the current extractor state.

protected override FixedWidthReport CreateProgressReport()

Returns

FixedWidthReport

A FixedWidthReport snapshot containing Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.CurrentItemCount, Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.CurrentSkippedItemCount, CurrentRejectedItemCount, CurrentFilteredLineCount, and CurrentLineNumber at the moment of the call.

CreateProgressTimer(IProgress<FixedWidthReport>)

Creates the Wolfgang.Etl.Abstractions.IProgressTimer used to drive progress callbacks. Override this method in a derived class to inject a custom timer (for example, a custom implementation that allows manual control in unit tests).

protected override IProgressTimer CreateProgressTimer(IProgress<FixedWidthReport> progress)

Parameters

progress IProgress<FixedWidthReport>

The progress sink that will receive callbacks.

Returns

IProgressTimer

A started Wolfgang.Etl.Abstractions.IProgressTimer instance.

Dispose(bool)

Releases the internal StreamReader when this instance was constructed from a Stream, then defers to the base class. Has no effect on a caller-owned TextReader.

protected override void Dispose(bool disposing)

Parameters

disposing bool

true when called from IDisposable; false when called from a finalizer.

ExtractWorkerAsync(CancellationToken)

This method is the core implementation of the extraction logic and should be overridden by derived classes.

protected override IAsyncEnumerable<TRecord> ExtractWorkerAsync(CancellationToken token)

Parameters

token CancellationToken

A CancellationToken to observe while waiting for the task to complete.

Returns

IAsyncEnumerable<TRecord>

IAsyncEnumerable<TSource> The result may be an empty sequence if no data is available or if the extraction fails.

OnItemError(ItemErrorContext)

Translates this extractor's MalformedLineHandling knob into the base per-item error policy. Skip maps to Wolfgang.Etl.Abstractions.ItemErrorAction.Skip; ThrowException maps to Wolfgang.Etl.Abstractions.ItemErrorAction.Abort. ReturnDefault recovers with a substitute record before the give-up decision, so it never reaches this hook.

protected override ItemErrorAction OnItemError(ItemErrorContext context)

Parameters

context ItemErrorContext

The failure context supplied by the extract worker.

Returns

ItemErrorAction

The action the base class should apply to the failed line.

Exceptions

InvalidOperationException

Thrown when MalformedLineHandling holds an unrecognized value.