Skip to content

About

Cudara is a self-hosted inference server for HuggingFace models. Run LLMs, Vision-Language Models, Embedding models, and Speech Recognition models on your GPU with an Ollama-compatible API.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages