Skip to content
@local-inference-lab

local-inference-lab

Popular repositories Loading

  1. rtx6kpro rtx6kpro Public

    RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

    Python 678 43

  2. sparkinfer sparkinfer Public

    Python 141 35

  3. llm-inference-bench llm-inference-bench Public

    LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.

    Python 59 8

  4. blackwell-llm-docker blackwell-llm-docker Public

    Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)

    Python 43 11

  5. vllm vllm Public

    Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 20 9

  6. quant-toolkit quant-toolkit Public

    Python 14 7

Repositories

Showing 10 of 12 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…