NousResearch/DeepHermes-Egregore-v1-RLAIF-8b-Atropos-GGUF Reinforcement Learning • 8B • Updated May 5 • 47 • 3
NousResearch/DeepHermes-AscensionMaze-RLAIF-8b-Atropos-GGUF Reinforcement Learning • 0.0B • Updated May 10 • 62 • 6
mradermacher/SEOcrate-4B_grpo_new_01-i1-GGUF Reinforcement Learning • 4B • Updated 18 days ago • 4.82k
Nellyw888/VeriReason-Qwen2.5-7b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • 8B • Updated May 31 • 871 • 4
vegeta03/DeepHermes-Egregore-v2-RLAIF-8b-Atropos-Q8_0-GGUF Reinforcement Learning • 8B • Updated May 9 • 9
vegeta03/DeepHermes-AscensionMaze-RLAIF-8b-Atropos-Q8_0-GGUF Reinforcement Learning • 8B • Updated May 9 • 6
vegeta03/DeepHermes-ToolCalling-Specialist-Atropos-Q8_0-GGUF Reinforcement Learning • 8B • Updated May 9 • 10
ajagota71/pythia-70m-detox-irl-rlhf-test-facebook-filter Reinforcement Learning • 0.1B • Updated May 11 • 30
Nellyw888/VeriReason-codeLlama-7b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • 7B • Updated May 31 • 835 • 2
Nellyw888/VeriReason-Qwen2.5-3b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • 3B • Updated May 31 • 27
Nellyw888/VeriReason-Qwen2.5-1.5b-RTLCoder-Verilog-GRPO-reasoning-tb Reinforcement Learning • 2B • Updated May 20 • 38 • 1
ajagota71/pythia-70m-fb-detox-checkpoint-epoch-60 Reinforcement Learning • 0.1B • Updated May 16 • 20
ajagota71/pythia-70m-fb-detox-checkpoint-epoch-100 Reinforcement Learning • 0.1B • Updated May 16 • 3