Llama 4 - a AmpereComputing Collection

AmpereComputing 's Collections

Gemma 3

GLM 4.5

GPT-OSS

Llama 4

Mistral

Phi 4

Qwen 3

QwQ

Llama 4

updated Sep 12

Ampere's quantization formats (Q4_K_4 / Q8R16) require Ampere optimized llama.cpp available here: https://hub.docker.com/r/amperecomputingai/llama.cpp

AmpereComputing/llama-4-scout-16e-17b-instruct-gguf

108B • Updated Jul 25 • 26