michaelbenayoun/llama-2-tiny-4kv-heads-4layers-random Text Generation • Updated 5 days ago • 11.5k
michaelbenayoun/llama-2-tiny-4kv-heads-16layers-random Text Generation • Updated 12 days ago • 8.68k
Running 2.66k 2.66k The Ultra-Scale Playbook 🌌 The ultimate guide to training LLM on large GPU Clusters