Class FixedWidthExtractor<TRecord>
- Namespace
- Wolfgang.Etl.FixedWidth
- Assembly
- Wolfgang.Etl.FixedWidth.dll
Reads a fixed-width text file and yields records of type TRecord
as an asynchronous stream.
public class FixedWidthExtractor<TRecord> : ExtractorBase<TRecord, FixedWidthReport>, IExtractWithProgressAndCancellationAsync<TRecord, FixedWidthReport>, IExtractWithCancellationAsync<TRecord>, IExtractWithProgressAsync<TRecord, FixedWidthReport>, IExtractAsync<TRecord>, IReportsItemErrors, IAsyncDisposable, IDisposable where TRecord : notnull, new()
Type Parameters
TRecordThe POCO type representing a single record. Properties decorated with FixedWidthFieldAttribute are populated from each line. The type must have a public parameterless constructor.
- Inheritance
-
ExtractorBase<TRecord, FixedWidthReport>FixedWidthExtractor<TRecord>
- Implements
-
IExtractWithProgressAndCancellationAsync<TRecord, FixedWidthReport>IExtractWithCancellationAsync<TRecord>IExtractWithProgressAsync<TRecord, FixedWidthReport>IExtractAsync<TRecord>IReportsItemErrors
- Inherited Members
-
ExtractorBase<TRecord, FixedWidthReport>.ExtractAsync()ExtractorBase<TRecord, FixedWidthReport>.CreateProgressReport()ExtractorBase<TRecord, FixedWidthReport>.IncrementCurrentItemCount()ExtractorBase<TRecord, FixedWidthReport>.IncrementCurrentSkippedItemCount()ExtractorBase<TRecord, FixedWidthReport>.OnItemError(ItemErrorContext)ExtractorBase<TRecord, FixedWidthReport>.HandleItemError(ItemErrorContext)ExtractorBase<TRecord, FixedWidthReport>.DisposeAsync()ExtractorBase<TRecord, FixedWidthReport>.Dispose()ExtractorBase<TRecord, FixedWidthReport>.StartedAtExtractorBase<TRecord, FixedWidthReport>.ElapsedExtractorBase<TRecord, FixedWidthReport>.ReportingIntervalExtractorBase<TRecord, FixedWidthReport>.CurrentItemCountExtractorBase<TRecord, FixedWidthReport>.CurrentSkippedItemCountExtractorBase<TRecord, FixedWidthReport>.CurrentErrorItemCountExtractorBase<TRecord, FixedWidthReport>.MaximumItemCountExtractorBase<TRecord, FixedWidthReport>.SkipItemCountExtractorBase<TRecord, FixedWidthReport>.WorkerResilienceExtractorBase<TRecord, FixedWidthReport>.ErrorPolicy
Examples
// Stream-based (preferred for files — 64 KB buffer reduces syscall overhead):
await using var stream = File.OpenRead("data.txt");
using var extractor = new FixedWidthExtractor<CustomerRecord>(stream);
// TextReader-based (caller owns the reader):
var extractor = new FixedWidthExtractor<CustomerRecord>(reader);
Remarks
Two construction modes are supported, each with different ownership semantics:
- TextReader constructor — the caller owns the TextReader lifetime. The extractor does not dispose it. Calling Dispose() is optional and has no effect.
- Stream constructor — the extractor creates an internal StreamReader with a 64 KB buffer for improved throughput on large files. The caller retains ownership of the Stream (it is not closed), but Dispose() must be called to release the internal reader.
Constructors
FixedWidthExtractor(Stream, ILogger<FixedWidthExtractor<TRecord>>, Encoding)
Initializes a new FixedWidthExtractor<TRecord> from a Stream decoded with the supplied Encoding, with diagnostic logging.
[Obsolete("Use the constructor that takes FixedWidthExtractorStreamOptions. This overload will be removed in a future release.")]
public FixedWidthExtractor(Stream stream, ILogger<FixedWidthExtractor<TRecord>> logger, Encoding encoding)
Parameters
streamStreamThe stream to use.
loggerILogger<FixedWidthExtractor<TRecord>>The logger to use for diagnostic output.
encodingEncodingThe encoding to decode with, or null for the documented default.
Exceptions
- ArgumentNullException
streamis null.
FixedWidthExtractor(Stream, Encoding)
Initializes a new FixedWidthExtractor<TRecord> from a Stream decoded with the supplied Encoding.
[Obsolete("Use the constructor that takes FixedWidthExtractorStreamOptions. This overload will be removed in a future release.")]
public FixedWidthExtractor(Stream stream, Encoding encoding)
Parameters
streamStreamThe stream to use.
encodingEncodingThe encoding to decode with, or null for the documented default.
Exceptions
- ArgumentNullException
streamis null.
FixedWidthExtractor(Stream, FixedWidthExtractorStreamOptions<TRecord>?, ILogger<FixedWidthExtractor<TRecord>>?)
Initializes a new FixedWidthExtractor<TRecord> over the specified Stream, with the logger as the trailing optional parameter.
public FixedWidthExtractor(Stream stream, FixedWidthExtractorStreamOptions<TRecord>? options = null, ILogger<FixedWidthExtractor<TRecord>>? logger = null)
Parameters
streamStreamThe Stream to use.
optionsFixedWidthExtractorStreamOptions<TRecord>Options that control behaviour, including the Encoding to use. When
null, the documented defaults apply.loggerILogger<FixedWidthExtractor<TRecord>>An optional logger for diagnostic output. When
null— or omitted — Instance is used and logging is disabled.
Exceptions
- ArgumentNullException
streamis null.
FixedWidthExtractor(TextReader, ILogger<FixedWidthExtractor<TRecord>>?)
Initializes a new FixedWidthExtractor<TRecord> that reads from the specified TextReader with diagnostic logging.
public FixedWidthExtractor(TextReader reader, ILogger<FixedWidthExtractor<TRecord>>? logger = null)
Parameters
readerTextReaderThe TextReader to read fixed-width records from.
loggerILogger<FixedWidthExtractor<TRecord>>The logger instance for diagnostic output.
Exceptions
- ArgumentNullException
readerorloggeris null.
FixedWidthExtractor(TextReader, FixedWidthExtractorOptions<TRecord>?, ILogger<FixedWidthExtractor<TRecord>>?)
Initializes a new instance that reads from reader with the given configuration.
[SuppressMessage("ApiDesign", "RS0026:Do not add multiple public overloads with optional parameters", Justification = "Shipped shape: the Stream overload keeps its optional options/logger for source compatibility until the constructor set is settled in the 2026-12-15 removal wave (#373 / #343).")]
public FixedWidthExtractor(TextReader reader, FixedWidthExtractorOptions<TRecord>? options = null, ILogger<FixedWidthExtractor<TRecord>>? logger = null)
Parameters
readerTextReaderThe reader supplying the fixed-width lines. The caller owns it.
optionsFixedWidthExtractorOptions<TRecord>The parsing configuration. null (the default) keeps every default. The reader has already decoded its bytes, so this is the base record without an
Encoding.loggerILogger<FixedWidthExtractor<TRecord>>An optional logger.
Exceptions
- ArgumentNullException
readeris null.
Properties
BlankLineHandling
Specifies what happens when a truly blank line (zero length) is encountered in the file. Evaluated before the skip budget and Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount.
public BlankLineHandling BlankLineHandling { get; set; }
Property Value
Remarks
- ThrowException (default) — always throws a LineTooShortException regardless of position.
- Skip — the line is invisible to all counting logic. Does not count toward Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.SkipItemCount or Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount.
- ReturnDefault — a default
TRecordinstance is yielded. Counts toward the skip budget if within Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.SkipItemCount, otherwise counts toward Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount.
Note: a line consisting entirely of spaces is not blank — it is a valid data line that will parse to a record with all whitespace-trimmed (empty/default) fields.
LineFilter is not invoked for blank lines.
CurrentByteOffset
The byte offset of the start of the next unread line — the value to persist as a checkpoint after processing each record. Advances as every physical line is consumed (including headers, blank, and skipped lines) so a resume continues exactly where reading stopped.
public long CurrentByteOffset { get; }
Property Value
Remarks
Thread-safe: read with Interlocked so it may be sampled from a progress timer thread.
Exceptions
- InvalidOperationException
Byte-offset tracking is not enabled — set TrackByteOffset or StartByteOffset first, and construct the extractor from a Stream.
CurrentFilteredLineCount
The number of physical lines read that did not produce a record and were
neither skipped by the budget nor rejected: header lines, the separator line,
blank lines dropped per BlankLineHandling, lines dropped by
LineFilter, and the line that triggered early termination. With
CurrentLineNumber this closes the line accounting:
CurrentLineNumber = CurrentItemCount + CurrentSkippedItemCount +
CurrentRejectedItemCount + CurrentFilteredLineCount.
public int CurrentFilteredLineCount { get; }
Property Value
CurrentLineNumber
The 1-based physical line number of the line most recently read from the file. Updated before each line is parsed so that if an exception is thrown, this value points to the offending line. Matches the line number shown in a text editor — no adjustment is needed for header or separator lines.
public long CurrentLineNumber { get; }
Property Value
Remarks
Thread-safe: reads are performed with Interlocked so this property may be sampled from a progress-reporting timer thread without a data race.
CurrentRejectedItemCount
The number of parsed records rejected so far: records discarded via
Skip and records rejected by
RecordValidator. Distinct from
Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.CurrentSkippedItemCount
(the SkipItemCount pagination budget).
public int CurrentRejectedItemCount { get; }
Property Value
FieldDelimiter
The delimiter string present between fields in the source file, or null (default) for pure fixed-width input with no delimiter. Must match the FieldDelimiter used when the file was written.
public string? FieldDelimiter { get; set; }
Property Value
Examples
// File was written with FieldDelimiter = " | " — set the same value on the extractor:
// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
FieldDelimiter = " | ",
// Pure fixed-width file with no delimiter (default):
FieldDelimiter = null,
Remarks
When set, the extractor accounts for the delimiter width when calculating field start positions, ensuring each field is read from the correct offset.
FieldSeparator
When non-null, the line immediately following the last header line is treated as a separator and skipped. Has no effect if HeaderLineCount is 0. The value of the character is not used for parsing — only its presence matters. Set to null (default) for no separator. Mirrors FieldSeparator.
public char? FieldSeparator { get; set; }
Property Value
- char?
Examples
// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
HeaderLineCount = 1,
FieldSeparator = '-', // skips a "----------" separator line after the header
FieldSeparator = null, // no separator line (default)
HasHeader
Convenience property — when set to true, sets HeaderLineCount to 1. When set to false, sets HeaderLineCount to 0. Returns true if HeaderLineCount is greater than zero. Mirrors WriteHeader.
public bool HasHeader { get; set; }
Property Value
Examples
// Skip one header line before reading records:
// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
HeaderLineCount = 1,
// Skip one header line and one separator line:
HeaderLineCount = 1,
FieldSeparator = '-',
HeaderLineCount
The number of header lines to skip at the beginning of the file before extracting records. Defaults to 0. For the common single-header case, use HasHeader instead.
public int HeaderLineCount { get; set; }
Property Value
Examples
// File has two header lines followed by data:
// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
HeaderLineCount = 2,
// Equivalent shorthand for the common single-header case:
HeaderLineCount = 1,
Remarks
Lines 1 through HeaderLineCount are skipped entirely without parsing. If FieldSeparator is also set, the line immediately after the last header line is additionally skipped as a separator line.
LineFilter
A delegate invoked for every data line (after header and separator lines have been skipped) before any parsing occurs. Return Process to parse the line normally, Skip to skip it, or Stop to end the stream immediately without parsing the line.
public Func<string, LineAction> LineFilter { get; set; }
Property Value
Examples
// Footer string — stop when a known marker line is reached
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => line == "END" ? LineAction.Stop : LineAction.Process }
// Trailing separator — stop when a line consists entirely of dashes
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => line.All(c => c == '-') ? LineAction.Stop : LineAction.Process }
// EOF marker — stop when a line starts with a sentinel prefix
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => line.StartsWith("$$") ? LineAction.Stop : LineAction.Process }
// Comment lines — skip lines that begin with '#'
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => line.StartsWith("#") ? LineAction.Skip : LineAction.Process }
// Blank line as terminator — stop at the first empty line
new FixedWidthExtractorOptions<TRecord> { LineFilter = line => string.IsNullOrWhiteSpace(line) ? LineAction.Stop : LineAction.Process }
Remarks
Evaluated after BlankLineHandling — blank lines never reach the filter.
Evaluated before the skip budget and Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount.
Both Skip and Stop are invisible
to all counting logic — they do not affect Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.SkipItemCount,
Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.MaximumItemCount, or
CurrentSkippedItemCount.
Defaults to a function that always returns Process.
MalformedLineHandling
Specifies what happens when a line is encountered that is too short or whose field values cannot be converted to the target property type.
public MalformedLineHandling MalformedLineHandling { get; set; }
Property Value
Remarks
Defaults to ThrowException.
When set to Skip, the line is skipped and
CurrentSkippedItemCount is incremented.
When set to ReturnDefault, a default instance
of TRecord is yielded for the offending line.
OnError
An optional dead-letter sink invoked once for each record that fails to parse (#29). The failed line is reported as a FixedWidthError — its 1-based source line number, raw content, and exception — so the caller can log it, push it to a dead-letter queue, or collect it for post-run inspection. Set MalformedLineHandling to Skip to capture-and-continue; with the default ThrowException the failure is still reported here before the exception is re-thrown. Business rejects from RecordValidator are not reported here — they are not parse errors.
public Action<FixedWidthError>? OnError { get; set; }
Property Value
Examples
var errors = new List<FixedWidthError>();
var extractor = new FixedWidthExtractor<Record>(reader, new FixedWidthExtractorOptions<Record>
{
MalformedLineHandling = MalformedLineHandling.Skip,
OnError = errors.Add,
});
await foreach (var ok in extractor.ExtractAsync(token)) { /* only good records */ }
// errors now holds the dead letters
RecordValidator
An optional callback invoked for each fully parsed record, after
ValueParser and field assignment but before the record is
yielded. Return Accept() to yield it,
Skip(string?) to drop it (increments
CurrentSkippedItemCount), or Stop(string?) to
end extraction. null (the default) applies no validation.
public Func<TRecord, ValidationResult>? RecordValidator { get; set; }
Property Value
- Func<TRecord, ValidationResult>
Examples
// Supplied through FixedWidthExtractorOptions<TRecord>, passed to the constructor:
RecordValidator = record =>
record.Balance < 0
? ValidationResult.Skip("Negative balance")
: ValidationResult.Accept(),
Schema
An optional layout that overrides the [FixedWidthField] / [FixedWidthSkip] attributes
on TRecord (#23). Build one with FixedWidthSchemaBuilder<T> to
map a type you cannot decorate, or to define the layout in code. When null (the
default) the attribute-based layout is used. The schema's RecordType
must be TRecord.
public FixedWidthSchema? Schema { get; set; }
Property Value
StartByteOffset
The byte offset to seek to before extraction begins — a checkpoint saved from a prior run's CurrentByteOffset. Defaults to 0 (start of stream). A non-zero value requires a seekable Stream-constructed extractor and implicitly enables TrackByteOffset. On resume, header lines are not re-skipped (the checkpoint is past them); Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.SkipItemCount is applied from the resumed position.
public long StartByteOffset { get; set; }
Property Value
Exceptions
- ArgumentOutOfRangeException
The value is negative.
TrackByteOffset
Enables byte-offset tracking for checkpoint/resume (#31). When true, CurrentByteOffset reports the byte position of the next unread line so a pipeline can save a checkpoint after each record and resume via StartByteOffset after a crash. Defaults to false — tracking is opt-in because it wraps the reader in a byte-counting decoder, changing the read path. Requires the Stream constructor; a caller-owned TextReader has no addressable byte stream. Setting StartByteOffset to a non-zero value enables tracking implicitly.
public bool TrackByteOffset { get; set; }
Property Value
Remarks
Byte offsets are computed with the encoding passed to the Stream constructor (UTF8 by default), so tracking assumes the stream is actually encoded that way. A byte-order mark that would switch the internal StreamReader to a different encoding (for example a UTF-16 BOM when UTF-8 was specified) is not supported for tracking and would yield incorrect offsets — pass the stream's real encoding to the constructor when enabling tracking. A matching-encoding BOM (e.g. a UTF-8 BOM with the default encoding) is handled correctly.
ValueParser
A delegate that converts a raw string read from the file into the target property type. The FieldContext provides the property type, format string, and other field metadata needed to perform the conversion. Defaults to DefaultParser. Mirrors ValueConverter.
public FixedWidthValueParser ValueParser { get; set; }
Property Value
Examples
// Treat "Y"/"N" as bool, fall back to DefaultParser for everything else:
new FixedWidthExtractorOptions<TRecord>
{
ValueParser = (text, ctx) =>
ctx.PropertyType == typeof(bool)
? (object)(text.Span.SequenceEqual("Y".AsSpan()))
: FixedWidthConverter.DefaultParser(text, ctx),
}
// Parse a custom date format for a specific field:
new FixedWidthExtractorOptions<TRecord>
{
ValueParser = (text, ctx) =>
ctx.PropertyName == "BirthDate"
? DateTime.ParseExact(text.ToString(), "dd/MM/yyyy", CultureInfo.InvariantCulture)
: FixedWidthConverter.DefaultParser(text, ctx),
}
Remarks
The delegate must return a value that is assignable to the property's CLR type. Returning an incompatible type will cause an InvalidCastException when the framework attempts to set the property value.
Exceptions
- FieldConversionException
The default DefaultParser wraps any parse failure in a FieldConversionException. Custom parsers should do the same so that MalformedLineHandling can handle them uniformly.
Methods
CreateProgressReport()
Creates a progress report snapshot for the current extractor state.
protected override FixedWidthReport CreateProgressReport()
Returns
- FixedWidthReport
A FixedWidthReport snapshot containing Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.CurrentItemCount, Wolfgang.Etl.Abstractions.ExtractorBase<TSource, TProgress>.CurrentSkippedItemCount, CurrentRejectedItemCount, CurrentFilteredLineCount, and CurrentLineNumber at the moment of the call.
CreateProgressTimer(IProgress<FixedWidthReport>)
Creates the Wolfgang.Etl.Abstractions.IProgressTimer used to drive progress callbacks. Override this method in a derived class to inject a custom timer (for example, a custom implementation that allows manual control in unit tests).
protected override IProgressTimer CreateProgressTimer(IProgress<FixedWidthReport> progress)
Parameters
progressIProgress<FixedWidthReport>The progress sink that will receive callbacks.
Returns
- IProgressTimer
A started Wolfgang.Etl.Abstractions.IProgressTimer instance.
Dispose(bool)
Releases the internal StreamReader when this instance was constructed from a Stream, then defers to the base class. Has no effect on a caller-owned TextReader.
protected override void Dispose(bool disposing)
Parameters
disposingbooltrue when called from IDisposable; false when called from a finalizer.
ExtractWorkerAsync(CancellationToken)
This method is the core implementation of the extraction logic and should be overridden by derived classes.
protected override IAsyncEnumerable<TRecord> ExtractWorkerAsync(CancellationToken token)
Parameters
tokenCancellationTokenA CancellationToken to observe while waiting for the task to complete.
Returns
- IAsyncEnumerable<TRecord>
IAsyncEnumerable<TSource> The result may be an empty sequence if no data is available or if the extraction fails.
OnItemError(ItemErrorContext)
Translates this extractor's MalformedLineHandling knob into the base per-item error policy. Skip maps to Wolfgang.Etl.Abstractions.ItemErrorAction.Skip; ThrowException maps to Wolfgang.Etl.Abstractions.ItemErrorAction.Abort. ReturnDefault recovers with a substitute record before the give-up decision, so it never reaches this hook.
protected override ItemErrorAction OnItemError(ItemErrorContext context)
Parameters
contextItemErrorContextThe failure context supplied by the extract worker.
Returns
- ItemErrorAction
The action the base class should apply to the failed line.
Exceptions
- InvalidOperationException
Thrown when MalformedLineHandling holds an unrecognized value.