Skip to content

feat: streaming and large-data support (#18) - #19

Merged
adiled merged 1 commit into
mainfrom
dev
Aug 27, 2026
Merged

adiled merged 1 commit into
mainfrom
dev

Conversation

@adiled

@adiled adiled commented Aug 27, 2026

Copy link
Copy Markdown
Owner
  • feat: streaming and large-data support
  • Existing functions accept sync iterables/generators (+ Float64Array for OHLCV)
  • Add async variants resampleOhlcvAsync / resampleTicksByTimeAsync / resampleTicksByCountAsync with out-of-order healing (outOfOrderMs)
  • CLI: stream CSV/JSONL/NDJSON files line-by-line, incremental output, new jsonl format
  • Changeset (minor), README docs. Closes Add support for streams and large data formats #9.
  • feat: parquet file input (async only) and ESM-only package

  • chore: keep require() working via CJS wrapper; downgrade to minor

  • Add per-record map option; drop parquet alias guessing

  • New src/map.ts: OhlcvMap/TickMap (Record form: canonical field -> source key; function form: (record) => IOHLCV/TradeTick) + mapToOhlcv/mapToTick + async-iterable wrappers. Exported from index.
  • Library: resampleOhlcvAsync / resampleTicksByTimeAsync / resampleTicksByCountAsync gain options.map, applied to Parquet rows and object-shaped AsyncIterable items (tuples/Float64Array unaffected).
  • Parquet reader now reads exact canonical columns only; any other layout requires map. Removes the OHLCV/TICK alias tables and resolveColumn guessing (per user: aliasing can go now that map exists).
  • CLI: --map field=sourceKey flag (Record form) threads into CSV header resolution, JSON/JSONL object key remap, JSON array buffer path, and the Parquet branch. With --map a CSV first line is always the header.
  • Tests: parquet custom-name Record/function map + partial-map error + alias fixture now requires map; library CCXT-style Record/function map on object streams + tuple pass-through; CLI --map on CSV/JSON/parquet/pipe + invalid field. Fixtures ohlcv_custom.parquet, ticks_custom.parquet added.
  • README + changeset updated: map option (Record + function), CLI --map, aliasing removed.
  • chore: exclude .js.map sourcemaps from published package (16 files, 23.8kB)

  • perf(deps): drop lodash dependency; implement 7 helpers natively

Install footprint: 461 kB -> ~145 kB. The seven lodash helpers used (isPlainObject, sum, max, min, groupBy, sortBy, chunk) were the entire 315 kB lodash package as a runtime dep; each is now a native inline implementation (stable Array.sort, plain-object grouping, loop min/max/sum, slice-based chunking). Runtime deps are now commander + hyparquet only — we install only what we build.

  • perf(deps): replace commander with mri (tiny 4kB arg parser); fix mri config-mutation bug; clean unknown-option errors

* feat: streaming and large-data support

- Existing functions accept sync iterables/generators (+ Float64Array for OHLCV)
- Add async variants resampleOhlcvAsync / resampleTicksByTimeAsync / resampleTicksByCountAsync with out-of-order healing (outOfOrderMs)
- CLI: stream CSV/JSONL/NDJSON files line-by-line, incremental output, new jsonl format
- Changeset (minor), README docs. Closes #9.

* feat: parquet file input (async only) and ESM-only package

* chore: keep require() working via CJS wrapper; downgrade to minor

* Add per-record map option; drop parquet alias guessing

- New src/map.ts: OhlcvMap/TickMap (Record form: canonical field -> source
  key; function form: (record) => IOHLCV/TradeTick) + mapToOhlcv/mapToTick +
  async-iterable wrappers. Exported from index.
- Library: resampleOhlcvAsync / resampleTicksByTimeAsync /
  resampleTicksByCountAsync gain options.map, applied to Parquet rows and
  object-shaped AsyncIterable items (tuples/Float64Array unaffected).
- Parquet reader now reads exact canonical columns only; any other layout
  requires map. Removes the OHLCV/TICK alias tables and resolveColumn
  guessing (per user: aliasing can go now that map exists).
- CLI: --map field=sourceKey flag (Record form) threads into CSV header
  resolution, JSON/JSONL object key remap, JSON array buffer path, and the
  Parquet branch. With --map a CSV first line is always the header.
- Tests: parquet custom-name Record/function map + partial-map error + alias
  fixture now requires map; library CCXT-style Record/function map on object
  streams + tuple pass-through; CLI --map on CSV/JSON/parquet/pipe + invalid
  field. Fixtures ohlcv_custom.parquet, ticks_custom.parquet added.
- README + changeset updated: map option (Record + function), CLI --map,
  aliasing removed.

* chore: exclude .js.map sourcemaps from published package (16 files, 23.8kB)

* perf(deps): drop lodash dependency; implement 7 helpers natively

Install footprint: 461 kB -> ~145 kB. The seven lodash helpers used
(isPlainObject, sum, max, min, groupBy, sortBy, chunk) were the entire
315 kB lodash package as a runtime dep; each is now a native inline
implementation (stable Array.sort, plain-object grouping, loop min/max/sum,
slice-based chunking). Runtime deps are now commander + hyparquet only —
we install only what we build.

* perf(deps): replace commander with mri (tiny 4kB arg parser); fix mri config-mutation bug; clean unknown-option errors
@adiled
adiled merged commit cff3a71 into main Aug 27, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add support for streams and large data formats

1 participant