Pull down to go back
Trained a 3B Llama with GRPO on a chemistry RL environment. Drug-like molecules in 6 hours of GPU time. [P]

Trained a 3B Llama with GRPO on a chemistry RL environment. Drug-like molecules in 6 hours of GPU time. [P]

What if AI could design a new drug in 30 seconds? That's the future I'm betting on. Maybe 5 years out. Maybe 15. But it's coming. This week I built a tiny piece of it. A reinforcement learning environment where an AI agent designs drug molecules atom by atom. Add a fragment. Swap an atom. Build a scaffold. The environment scores it on real chemistry: Lipinski rules, drug-likeness, synthesis difficulty, target protein binding. Trained Llama-3.2-3B with GRPO. Six hours on a single A10G. Image