source: Hugging Face Blog: Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
level: technical
asyncgrpotrainer now supports lora, syncing only a small adapter to vllm instead of full weights. this lets training and inference run as separate hugging face jobs on different machines. a storage bucket mounted in every job carries the adapter files, removing the need for nccl or shared localhost. a proxy adds auth headers, routes rollouts by kv prefix, and broadcasts adapter loads to all replicas. this setup turns a recipe that took 3 hours 27 minutes into 53 minutes for 500 steps.
the adapter is rank 1, just a few megabytes for a 1.5b model, versus 3 gb for full weights. vllm keeps multiple adapters loaded, so old rollouts finish with their starting policy while new ones use the latest. the trainer publishes a new adapter every four optimizer steps, with max_staleness=4 requiring six adapter slots. versioned adapter names prevent kv cache mismatches. the proxy tracks 16-token block hashes chained with the adapter name, so requests with shared prefixes hit the same replica and skip redundant prefill.
the dataset is sail/sanity-test-r1d-1.5b, 1,460 math problems with 20-80% success rates, giving clear early signal. hyperparameters follow the precision-rl paper: qwen2.5-math-1.5b, lora rank 1 alpha 2, learning rate 4e-5, 8 samples per prompt, 128 completions per step. vllm is pinned to v0.27.1 because flags and endpoints change fast. checkpoints and final adapters persist in the bucket, so a preempted trainer job can resume without losing work.
why it matters: this shows how to scale rl fine-tuning across separate cloud jobs without high-speed interconnects, making large-scale training cheaper and more accessible.
source: Hugging Face Blog: Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL