level: business
zml, a paris-based ai startup backed by yann lecun, released zml/llmd, a free inference server that lets open-source large language models run on chips from nvidia, amd, google, apple, and intel. founder steeve morin said the goal is to break silos and let different hardware work at peak speed for ai tasks. the software targets the growing need for efficient inference as ai use spreads, addressing patchy behind-the-scenes integration that often forces vendor lock-in.
the product could help enterprises and clouds mix cheaper or less power-hungry chips, easing fears about ai costs. morin noted it may also aid novel european chipmakers like axelera and sipearl. while not open source, llmd is free to gather usage data before any paid launch. zml competes with inference-focused firms like baseten and tools like vllm and sglang, but morin says his 20-person team is already co-designing silicon, hinting at broader ambitions.
zml raised $20 million from investors including 20vc and kima ventures, with angel backing from hugging face founders and docker's solomon hykes. morin credits the lean team and paris location for fast progress. the launch reflects europe's growing ability to build ai startups locally, with zml aiming to give users control over their ai systems and achieve real efficiency gains.
why it matters: it offers a practical way to run ai models on diverse hardware, potentially lowering costs and energy use while reducing dependence on single chip suppliers.