spec-driven agentic development reshapes software delivery
a new report formalizes spec-driven agentic development, where machine-readable specs fuel autonomous coding agents across the software lifecycle.
researchsummaries filed under ai
a new report formalizes spec-driven agentic development, where machine-readable specs fuel autonomous coding agents across the software lifecycle.
researchgoogle deepmind and fenris creations are collaborating to build ai agents that can learn, remember, and plan in the persistent world of eve online.
researchlondon startup inherent says its small ai agent faraday outperformed anthropic and openai models at reproducing scientific paper findings.
industryA new GPU scheduler from Dharma-AI improves cluster utilization and priority-weighted output by planning allocations across the entire scheduling horizon instead of using FIFO order.
machine learningNew probes show leading ASR models reproduce benchmark transcripts even when audio contradicts them, inflating scores.
machine learningGoogle Research introduces ME-POIs, a framework that blends text and mobility data to improve predictions about places like opening hours and busyness.
researchGoogle Research introduces a multi-agent AI system that structures biomarker discovery from wearable data through iterative hypothesis generation, statistical analysis, and literature-grounded reasoning.
researchIBM Research shows that the right amount of agentic memory varies by model tier, with curated retrieval boosting weaker models and full guideline sets helping stronger ones.
machine learningLiquid AI published DSpark draft checkpoints for three LFM2.5 models, enabling speculative decoding that speeds up inference up to 3.2x on GPU and 2.87x on-device without changing output quality.
machine learningAn analysis of 500 Hugging Face model cards finds current transparency artifacts insufficient for downstream safety governance of open-weight foundation models.
researchLiquid AI releases quantization-aware distillation checkpoints that recover 97% of BF16 accuracy for LFM2.5 models at Q4_0 memory and speed.
machine learningMojo, a Python-inspired language for GPU programming, has released its compiler and toolchain under Apache 2 license.
toolsSentence Transformers v6.0 introduces MultiVectorEncoder for ColBERT-style late interaction retrieval, supporting PyLate, Stanford-NLP ColBERT, and ColPali checkpoints.
machine learningGoogle Research demonstrates a deep learning method that predicts insulin resistance from smartphone imagery with accuracy close to DXA scans.
researchRubricForge evolves a human-readable rubric from labeled trajectories to reduce over-crediting in language-model agent evaluation.
researchA one-year production trace from Chutes reveals how LLM serving workloads evolve, with implications for caching and load-balancing.
researchThis week in AI: watermarking for Claude, open model shifts, reasoning debates, and new benchmarks for integrity and factuality.
data scienceAnthropic explains its text watermarking for Claude, using Google DeepMind's SynthID-Text to comply with the EU AI Act, and addresses editing, code, and detection.
industryAMD and Meta engineers upstreamed FP8 training optimizations for AMD Instinct GPUs into TorchAO and TorchTitan, delivering up to 13.4% throughput gains on dense models and recovering 89% of quantization overhead on MoE models.
machine learningA position paper argues that AI reasoning lacks clear definitions and proposes treating it as a learnable rule-based process with a best-practices checklist.
research