Skip to content

Repository files navigation

Tensora

Checkpoint loading for large language models.

Tensora is an open-source framework for loading LLM checkpoints. It supports SafeTensors and ServerlessLLM storage layouts with caller-selected I/O backends: synchronous POSIX, Tokio async, Linux io_uring, and memory-mapped lazy access.

Paper: Load by Design: Adaptive Heuristics for LLM Checkpoint Loading — Botir Khaltaev (2026). Sources in paper/.


Key Results

Regime Winner Mechanism
Small/single-shard SafeTensors sync Thread-parallel chunked POSIX reads
Large multi-shard SafeTensors (≥ 4 GB) io_uring Multi-worker ring submission
Range-heavy ServerlessLLM async Tokio grouped per-file tasks
Large partitioned ServerlessLLM io_uring Batched submission with coalescing

Use explicit backend flags to reproduce and compare these regimes.


Quick Start

# Build
cargo build --release

# Load a model (downloads from HuggingFace Hub on first run)
cargo run --release --bin profile -- safetensors default --model-id Qwen/Qwen3-0.6B --iterations 1

# Demo all I/O backends
cargo run --release --bin demo -- safetensors all --model-id Qwen/Qwen3-0.6B

Python Bindings

cd bindings/python
uv sync --group dev --group torch
uv run python examples/pytorch.py gpt2 --prompt "Hello"

Architecture

tensora/
├── src/
│   ├── storage/        # I/O backends (sync, tokio, io_uring, mmap)
│   ├── formats/        # Checkpoint formats (SafeTensors, ServerlessLLM)
│   ├── converters/     # Format conversion pipelines
│   ├── hf_model.rs     # HuggingFace Hub integration
│   └── bin/            # CLI tools (profile, demo, convert)
├── bindings/python/    # Python package (tensora_py) with PyTorch/vLLM support
├── benches/            # Criterion.rs benchmarks
├── scripts/            # Benchmark orchestration scripts
├── paper/              # LaTeX sources for the paper
└── results/            # Archived experiment data

Requirements

Component Version
Rust ≥ 1.92
OS Linux (full feature set with io_uring); other platforms lack io_uring
Python ≥ 3.12 (for bindings)
uv Latest (for Python workflows)

Testing

cargo test --lib --locked
cargo clippy --lib --locked -- -D warnings

Benchmarks

Rust (Criterion)

export TENSORA_MODEL_ID=openai-community/gpt2
cargo bench

Python (pytest-benchmark)

export TENSORA_BENCH_MODELS=openai-community/gpt2
./scripts/run_benchmarks.sh

Reproducing Paper Experiments

See the paper's Experimental Setup for cold-cache methodology. In brief:

# Drop caches (requires root)
sync && echo 3 > /proc/sys/vm/drop_caches

# Run profiling
cargo run --release --bin profile -- safetensors sync --model-id Qwen/Qwen3-8B --iterations 1

Replication targets storage-engine ordering and regime behaviour (e.g., SafeTensors crossover point, ServerlessLLM async advantage), not identical millisecond timings across hardware.


Citation

@software{khaltaev2026tensora,
  title   = {Load by Design: Adaptive Heuristics for LLM Checkpoint Loading},
  author  = {Khaltaev, Botir},
  year    = {2026},
  url     = {https://github.com/botirk38/tensora},
  license = {Apache-2.0}
}

See also: CITATION.cff and CITATION.bib.


License

Apache 2.0 — see LICENSE.

About

Adaptive checkpoint loading for large language models.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages