hf cli rebuilt for coding agents
hugging face redesigned its hf command-line tool to serve both humans and ai coding agents, cutting token usage by up to 6x on complex tasks.
topic
hugging face redesigned its hf command-line tool to serve both humans and ai coding agents, cutting token usage by up to 6x on complex tasks.
a guide to fine-tuning nvidia's multilingual streaming speech recognition model on custom data for better accuracy in under-resourced languages and domains.
new analysis of alternating power iteration for spiked tensor models gives finite-iteration error bounds and explains warm-start behavior without relying on specific initializations.
the ieee p3109 draft standard defines parameterized binary floating-point formats and operations to support efficient, consistent machine learning computations.
jetbrains launches mellum2, an open 12b-parameter mixture-of-experts model optimized for low-latency text and code tasks, activating only 2.5b parameters per token.
a large-scale analysis across four eeg datasets examines how different scalp regions contribute to predicting cognitive workload, revealing consistent patterns and practical implications for sensor selection.
a guide to five foundational papers covering transformer architecture, few-shot learning, scaling laws, instruction tuning, and retrieval-augmented generation for understanding large language models.
reachy mini robot can now use remote tools like weather and web search hosted on hugging face spaces via mcp, without downloading code locally.
deepspeed now integrates muon optimizer, reducing memory and improving convergence for large model training.
direct preference optimization reduced text degeneration by an average of 59.4% across five ocr model families by using the model's own failure outputs as rejection pairs.
a new framework lets human agents approve or reject algorithmic price suggestions, using old pricing data to skip the slow start typical in sparse booking markets.
a new method reduces the cubic complexity of gaussian processes with gradients by using exact gradient reduction and vecchia approximation.
periodic and soft target updates can guarantee convergence in linear q-learning under explicit spectral and step-size conditions.
a new method uses gradient tests instead of validation loss to decide when to stop training gradient boosted trees, avoiding the need for a patience parameter.
a new diagnostic reveals that common anomaly detection benchmarks become unreliable when held-out classes overlap with normal data in representation space.
holo3.1 expands computer-use ai to mobile, desktop, and web with quantized models for local inference.
an overview of advances in making large language models more interpretable through dynamic evaluation, statistical methods, and accessible tools.
linkedin rebuilt its dualip solver in pytorch to handle linear programs with trillions of variables, achieving 75x faster per-iteration times on gpus.