local-inference-lab
Popular repositories Loading
-
-
llm-inference-bench
llm-inference-bench PublicLLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
-
blackwell-llm-docker
blackwell-llm-docker PublicDocker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)
-
vllm
vllm PublicForked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
-
Repositories
- rtx6kpro Public
RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink
- blackwell-llm-docker Public
Docker images for LLM inference (SGLang + vLLM) on NVIDIA Blackwell GPUs (SM120, CUDA 13.2)
- sparkinfer Public
- vllm Public Forked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- fanpilot Public
Temperature/power-driven fan control for GPU servers — ASRock Rack BMC (case fans) + NVIDIA NVML (GPU fans), one daemon. Prometheus exporter + Grafana dashboard.
- llm-inference-bench Public
LLM inference decode throughput benchmark with Rich TUI dashboard. Measures token generation speed across concurrency levels and context lengths. Supports SGLang and vLLM engines.
- llmconduit Public
- quant-toolkit Public
- sglang Public Forked from sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
- nccl-canonical Public Forked from NVIDIA/nccl
Optimized primitives for collective multi-GPU communication
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…