Skip to content

feat: torch-free dataloader utilities for CPU-only tooling - #465

Merged
rrutmann merged 1 commit into
mainfrom
feat/weighted-combined-dataset
Oct 9, 2026
Merged

rrutmann merged 1 commit into
mainfrom
feat/weighted-combined-dataset

Conversation

@rrutmann

@rrutmann rrutmann commented Oct 9, 2026

Copy link
Copy Markdown
Collaborator

What does this PR do?

This PR guards the torch import in logger_utils so IndexGenerator and LargeFileLinesReader can be imported without torch, which modalities declares only in its cpu/cu12x extras. CPU-only preprocessing tooling no longer installs torch to call them.

Also adds TokenizerInstantiationModel so a tool can reuse a packing config for just its tokenizer.

General Changes

  • logger_utils.get_logger now imports torch defensively; when torch isn't installed, the rank prefix is simply omitted instead of raising ImportError. This unblocks importing create_index.IndexGenerator and large_file_lines_reader.LargeFileLinesReader — both used by CPU-only data-preprocessing tooling — without needing a torch install at all. Behavior is unchanged when torch is present.
  • TokenizerInstantiationModel (config/instantiation_models.py): lets a tool build just the tokenizer out of an existing packing config, without needing to satisfy the rest of that config's settings (e.g. a source file path it has no interest in).

Breaking Changes

  • None

Checklist before submitting final PR

  • My PR is minimal and addresses one issue in isolation
  • I have merged the latest version of the target branch into this feature branch
  • I have reviewed my own code w.r.t. correct implementation, missing type hints, proper documentation, etc.
  • I have run a sample config for model training
  • I have checked that all tests run through (python tests/tests.py)
  • I have updated the internal changelog (CHANGELOG_DEV.md)

Guards the torch import in logger_utils so IndexGenerator and LargeFileLinesReader
can be imported without torch, which modalities declares only in its cpu/cu12x
extras. CPU-only preprocessing tooling no longer installs torch to call them.

Adds TokenizerInstantiationModel so a tool can reuse a packing config for just
its tokenizer, without also having to satisfy that config's other settings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@rrutmann
rrutmann requested a review from le1nux October 9, 2026 14:35
@rrutmann rrutmann self-assigned this Oct 9, 2026
@rrutmann
rrutmann merged commit c4217d7 into main Oct 9, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants