source: PyTorch Blog: How Shopify built a continual learning loop with PyTorch and vLLM
level: technical
shopify built a continual learning loop that compresses production failures into model weights every day. the system starts with a frontier model, then uses a calibrated judge to score conversations. low-scoring examples are mined from anonymized traffic and repaired by a panel of frontier reasoning models. successful repairs become training data for supervised fine-tuning and reinforcement learning. this loop runs daily, allowing a smaller model to surpass the frontier baseline in quality.
the graphql agent serves up to 2,000 requests per minute. serving on a frontier model would cost an estimated $27 million per year, but the fine-tuned model costs closer to $1 million, a 96% reduction. gist compression reduced the system prompt from 6,000 tokens to 1,500 learned gist tokens. in load tests at 350 requests per minute, time-to-first-token dropped 19% and end-to-end latency dropped 38%. throughput rose 16% in requests per second, requiring 14% fewer gpus.
the loop relies on a rubric that defines quality criteria like completeness and safety. annotators label random traffic, and inter-annotator agreement is measured with cohen's kappa. the judge is calibrated using dspy with reflection-based optimizers. training uses pytorch for distributed full-parameter fine-tuning and vllm for serving. the approach shows how production experience can be turned into better weights, not just better prompts.
why it matters: this method lets companies run smaller, cheaper models that learn from real user interactions, reducing costs and latency while improving quality.
source: PyTorch Blog: How Shopify built a continual learning loop with PyTorch and vLLM