Providers · Local runtime
Hugging Face
CSuite runs ONNX builds of open models from Hugging Face fully locally, in-process via transformers.js. Like Ollama models, they are downloaded once and run on your own hardware — no API key and no per-use fees — with smaller models that work well even on modest machines.
Text(6 models)
Qwen3 4B
Alibaba's compact 4B instruction model with strong reasoning and multilingual capabilities.
1 year ago (Apr 2025)
Qwen3 0.6B
Alibaba's ultra-compact 0.6B Qwen3 model — runs in-process via HuggingFace or through Ollama, with tool use.
1 year ago (Apr 2025)
Gemma 3 1B
Google's compact 1B Gemma 3 model — runs in-process via HuggingFace or through Ollama.
1 year ago (Mar 2025)
R1 1.5B
DeepSeek's compact 1.5B reasoning model distilled from Qwen2.5 — chain-of-thought on-device.
1 year ago (Jan 2025)
Llama 3.2 1B
Meta's 1B instruction-tuned model — the smallest Llama, runnable in-process via HuggingFace or through Ollama.
2 years ago (Sep 2024)
Llama 3.2 3B
Meta's compact 3B instruction-tuned model — fast on-device text generation.
2 years ago (Sep 2024)