today's ai news covers hardware competition, new tools for developers, and research that questions common assumptions. amd unveiled a rack-scale system to rival nvidia, while a startup's valuation doubled on inference chip demand. on the software side, a model router promises to simplify generative ai workflows, and a dataset reveals real-world serving patterns.

  1. amd helios ai rack targets nvidia dominance - amd's new rack-scale system has major cloud and lab customers, signaling a serious challenge to nvidia's data center stronghold.
  2. etched hits $10.3b valuation for ai inference chips - the startup raised $300m and doubled its valuation in seven months, with chips in testing and $1b in orders, showing strong demand for specialized inference hardware.
  3. runway launches media router for generative ai models - a new router automatically picks the best image, video, or audio model based on quality, speed, or cost, simplifying multi-model workflows.
  4. fineserve dataset reveals real llm serving patterns - a dataset from a global platform exposes how real-world llm workloads vary, challenging assumptions from synthetic traces and informing better serving strategies.
  5. helion brings easier tpu kernel coding to pytorch - pytorch's helion dsl compiles to tpu code via pallas, letting users write performance-portable kernels without deep hardware expertise.

other notable stories include nvidia gpus heading to the moon on a lunar rover, a free open-source voice ai app that runs locally, and pypi blocking new file uploads to old releases to prevent supply chain attacks. research also explored bayesian model selection in transformers and molecular design with decoupled annealing flows.