Micro Models
9 models in your fleet
GPT-OSS-120B
DeployedOpen-weight frontier-scale language model for advanced reasoning, coding, and long-context enterprise workloads
LanguageCode
Feb 18, 2026120B paramsv1.0.0
GPT-OSS-20B
ReadyCompact open-weight model tuned for fast inference and on-device deployment
LanguageCode
Mar 11, 202620B paramsv1.0.0
Qwen2.5
ReadyMultilingual model with strong instruction-following, coding, and mathematical reasoning
LanguageCodeMath
Mar 5, 202672B paramsv2.5.0
Qwen3
ReadyNext-generation multimodal model unifying language and vision understanding
LanguageImage
Mar 12, 202632B paramsv3.0.0-alpha
Qwen3.5
QueuedAudio-augmented conversational model for speech-aware assistants
AudioLanguage
Mar 13, 202614B paramsv3.5.0
Llama3
DeployedGeneral-purpose open-weight model balancing language, code, and reasoning tasks
MathCodeLanguage
Jan 22, 202670B paramsv3.0.0
Qwen3.6
ReadyRefined iteration of Qwen3 with improved coding accuracy and tool-use reliability
LanguageCode
Mar 8, 202635B paramsv3.6.0
Gemma4
ReadyLightweight multimodal model optimized for on-device and edge inference
LanguageImage
Mar 10, 202627B paramsv4.0.0-beta
Gemma2
QueuedCompact, efficient language model for low-latency conversational applications
Language
Mar 13, 20269B paramsv2.0.0