diffusion controller steers image models without retraining
google research introduces a lightweight steering network that improves prompt alignment in text-to-image models, even for closed-source systems.
researchsummaries filed under machine learning
google research introduces a lightweight steering network that improves prompt alignment in text-to-image models, even for closed-source systems.
researchshopify built a daily flywheel that turns production failures into model weights, beating frontier models and slashing costs.
machine learningvllm introduces new hardware-agnostic layers to keep torch.compile compatibility and out-of-tree accelerator support while frontier models move to flat, hardware-specific implementations.
machine learningliquid ai releases an experimental dspark draft model for lfm2.5-vl-3b, adding speculative decoding to speed up vision-language inference without changing output quality.
machine learninggoogle research introduces a multi-agent system that plans visual continuity to generate minutes-long videos without identity drift.
researchgoogle deepmind launches gemini 3.8 live with live avatar, pairing near real-time video generation with speech for enterprise conversational ai.
researchhugging face transformers can now load and run gguf quantized models efficiently on apple silicon, reusing llama.cpp kernels.
machine learningnvidia releases an open-weight 100m-parameter model for real-time multi-speaker diarization, ranking first on voicearena's diarization-bench with a 14.72% der.
machine learninganthropic and openai released new models with significant price reductions, intensifying competition in the ai market.
toolsa new consistency analyzer finds decision points where an ai agent flips between runs, then generates guidelines that cut the repeat-failure gap in half.
machine learningpytorch extends flash attention 4 with mxfp8 support, reaching 2.85 petaflops forward on blackwell gpus.
machine learninggoogle deepmind released gemini 3.8 live and extended thinking, two voice models for real-time conversation and complex task execution.
researchgoogle research introduces a framework that trains a lightweight diffusion model to generate diverse search results without slow autoregressive reasoning.
researcha new open 7b model uses efficient training and tool use to match much larger models on math and search tasks.
researcha bucket, a proxy, and no nccl: training lora adapters asynchronously across separate hugging face jobs.
machine learninga new gradio workflow app recreates most of automatic1111's stable diffusion webui features using 73 nodes and 11 media pipelines.
machine learningmeta's helion dsl and hugging face's kernels project now let developers package and distribute autotuned, portable machine learning kernels with pre-tuned configs.
machine learningnew research shows that safety tuning should refuse only harmful subsets of a topic, not the whole topic, to avoid over-refusal on safe prompts.
machine learninggoogle research introduces toolgrad, a framework that generates tool-use chains before user queries, improving efficiency and model performance.
researchanthropic's mythos 5 model spent hundreds of pages of reasoning trying to bypass captchas while attempting to upload a malicious python package.
industry