granite 4.2: how ibm built its reasoning llms
ibm released granite 4.2, a family of dense reasoning llms in 3b, 8b, and 30b sizes, trained from scratch on 15t tokens with a multi-stage rl pipeline.
machine learningsummaries filed under machine learning
ibm released granite 4.2, a family of dense reasoning llms in 3b, 8b, and 30b sizes, trained from scratch on 15t tokens with a multi-stage rl pipeline.
machine learningkvboost cuts llm prefill latency by reusing key-value cache chunks at arbitrary prompt positions, not just shared prefixes.
researchgoogle deepmind and fenris creations are collaborating to build ai agents that can learn, remember, and plan in the persistent world of eve online.
researchA new GPU scheduler from Dharma-AI improves cluster utilization and priority-weighted output by planning allocations across the entire scheduling horizon instead of using FIFO order.
machine learningNew probes show leading ASR models reproduce benchmark transcripts even when audio contradicts them, inflating scores.
machine learningGoogle Research introduces ME-POIs, a framework that blends text and mobility data to improve predictions about places like opening hours and busyness.
researchGoogle Research introduces a multi-agent AI system that structures biomarker discovery from wearable data through iterative hypothesis generation, statistical analysis, and literature-grounded reasoning.
researchIBM Research shows that the right amount of agentic memory varies by model tier, with curated retrieval boosting weaker models and full guideline sets helping stronger ones.
machine learningLiquid AI published DSpark draft checkpoints for three LFM2.5 models, enabling speculative decoding that speeds up inference up to 3.2x on GPU and 2.87x on-device without changing output quality.
machine learningLiquid AI releases quantization-aware distillation checkpoints that recover 97% of BF16 accuracy for LFM2.5 models at Q4_0 memory and speed.
machine learningSentence Transformers v6.0 introduces MultiVectorEncoder for ColBERT-style late interaction retrieval, supporting PyLate, Stanford-NLP ColBERT, and ColPali checkpoints.
machine learningGoogle Research demonstrates a deep learning method that predicts insulin resistance from smartphone imagery with accuracy close to DXA scans.
researchAMD and Meta engineers upstreamed FP8 training optimizations for AMD Instinct GPUs into TorchAO and TorchTitan, delivering up to 13.4% throughput gains on dense models and recovering 89% of quantization overhead on MoE models.
machine learningHugging Face data from early 2026 shows Chinese labs dominating frontier open models, small models driving downloads, and coding agents becoming a major user base.
machine learningA community hackathon used coding agents to reproduce over 2,200 ICML 2026 papers, finding that 23% had at least one falsified or contested claim.
machine learningGoogle DeepMind introduces Gemini 3.7 Flash, a more intelligent and cost-effective model for coding and agent workflows.
researchGoogle DeepMind's SL2T model powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with ASL to English.
researchGoogle Research introduces knowledge profiling to show frontier LLMs encode nearly all facts but struggle to recall them, with thinking recovering many failures.
researchMeta open-sourced Muse Glimmer, a 30B-parameter model distilled from Muse Spark, optimized for on-device agentic workflows with ExecuTorch support for NVIDIA GPUs and Apple Silicon.
machine learningMeta’s Muse Glimmer is a 30B-parameter multimodal model optimized for local, privacy-aware agentic tasks like coding and document analysis, released under Apache 2.0 with day-0 support in transformers and llama.cpp.
machine learning