pytorch 2.13: flexattention on apple silicon, fused loss, new distributed backend
pytorch 2.13 brings flexattention to mps, a fused linear and cross-entropy loss, a new torchcomms backend, and more.
topic
pytorch 2.13 brings flexattention to mps, a fused linear and cross-entropy loss, a new torchcomms backend, and more.
a single lightweight controller coordinates attention mode, expert selection, and cache bit-width per token to reduce inference cost without losing quality.
the transformers modeling backend in vllm now achieves throughput equal to or better than hand-written native implementations for many large language models.
nvidia nemotron releases open datasets and a prompt atlas to help developers build inspectable, reliable ai agents using synthetic data.
a new mathematical model uses fiber bundles to detect when machine learning systems discover causal laws instead of interpolating data.
finer temporal disaggregation inflates in-sample fit but degrades out-of-sample accuracy due to recursive error compounding.
hugging face and amazon sagemaker ai now offer a deep-link integration that lets developers move from model discovery to fine-tuning or deployment in sagemaker studio with a single click.
sqlite-utils 4.0 introduces database schema migrations, nested transactions, and compound foreign keys for sqlite databases.
mount hugging face buckets and repos into skypilot jobs on any cloud with no egress fees for reads, using one hf:// url and your existing token.
a comparison of sql, pandas, and a claude agent on three analytics problems across speed, accuracy, and other dimensions.
microsoft foundry now offers a curated catalog of open-weight hugging face models deployable with one click onto managed compute, refreshed weekly and pre-staged in azure.
render pdf pages to images and use gemma 4 to extract structured data without ocr or layout parsers.
a plain-language guide to ten essential probability ideas that make machine learning work, from random variables to entropy.
a study shows that perturbation-based construct-validity audits can produce misleading conclusions due to hidden pipeline failures, and proposes a due-diligence gate to catch them.
a new method lets text-to-speech models pronounce tricky words correctly by using short audio hints, without changing the speaker's voice.
pytorch monarch now supports amd instinct gpus with rocm, enabling single-controller distributed training that recovers from node failures without full restarts.
photoroom details its data strategy for the prx text-to-image model, covering dataset assembly, captioning, and storage formats.
data scientists now spend more time managing ai systems than building models, with new roles in governance, prompt engineering, and agent supervision.