hugging face releases 207 webgpu kernels for faster browser ai
hugging face launched a library of 207 optimized webgpu kernels and a browser benchmarking tool to speed up local ai inference.
machine learningsummaries filed under tools
hugging face launched a library of 207 optimized webgpu kernels and a browser benchmarking tool to speed up local ai inference.
machine learninghugging face and voice arena release monsoon, a new open evaluation set for hindi and indian english speech recognition with detailed speaker metadata.
machine learninggoogle deepmind releases gemini omni 1.1 flash, a production-ready video model with scene extension, keyframe interpolation, 360p drafts, and 4k upscaling.
researchresearcher johann rehberger found a prompt injection attack that bypasses claude code's auto mode safety classifier 80% of the time.
toolsa new healing method lets a compressed, 4-bit model outperform its bfloat16 original on most benchmarks.
machine learningibm released granite 4.2, a family of dense reasoning llms in 3b, 8b, and 30b sizes, trained from scratch on 15t tokens with a multi-stage rl pipeline.
machine learninga system that spawns coding agents with pre-loaded memories from personal databases to avoid empty context windows.
researchA new GPU scheduler from Dharma-AI improves cluster utilization and priority-weighted output by planning allocations across the entire scheduling horizon instead of using FIFO order.
machine learningNew probes show leading ASR models reproduce benchmark transcripts even when audio contradicts them, inflating scores.
machine learningIBM Research shows that the right amount of agentic memory varies by model tier, with curated retrieval boosting weaker models and full guideline sets helping stronger ones.
machine learningLiquid AI published DSpark draft checkpoints for three LFM2.5 models, enabling speculative decoding that speeds up inference up to 3.2x on GPU and 2.87x on-device without changing output quality.
machine learningLiquid AI releases quantization-aware distillation checkpoints that recover 97% of BF16 accuracy for LFM2.5 models at Q4_0 memory and speed.
machine learningMojo, a Python-inspired language for GPU programming, has released its compiler and toolchain under Apache 2 license.
toolsSentence Transformers v6.0 introduces MultiVectorEncoder for ColBERT-style late interaction retrieval, supporting PyLate, Stanford-NLP ColBERT, and ColPali checkpoints.
machine learningAMD and Meta engineers upstreamed FP8 training optimizations for AMD Instinct GPUs into TorchAO and TorchTitan, delivering up to 13.4% throughput gains on dense models and recovering 89% of quantization overhead on MoE models.
machine learningHugging Face data from early 2026 shows Chinese labs dominating frontier open models, small models driving downloads, and coding agents becoming a major user base.
machine learningA community hackathon used coding agents to reproduce over 2,200 ICML 2026 papers, finding that 23% had at least one falsified or contested claim.
machine learningGoogle DeepMind introduces Gemini 3.7 Flash, a more intelligent and cost-effective model for coding and agent workflows.
researchMeta open-sourced Muse Glimmer, a 30B-parameter model distilled from Muse Spark, optimized for on-device agentic workflows with ExecuTorch support for NVIDIA GPUs and Apple Silicon.
machine learningMeta’s Muse Glimmer is a 30B-parameter multimodal model optimized for local, privacy-aware agentic tasks like coding and document analysis, released under Apache 2.0 with day-0 support in transformers and llama.cpp.
machine learning