Multiverse Computing, whose technology is compressing large language models to make them smaller, faster and cheaper, has raised $570m (€500m) in a Series C round at a $1.7bn valuation. The goal is to scale AI that can run on devices as small as smartphones, without relying on hyperscale cloud infrastructure.

CompactifAI is Multiverse’s compression technology that applies tensor networks (a quantum physics framework) to reduce the size of LLMs by up to 80-95% with “immaterial accuracy loss”, according to the company.

The framework that reduces the size of the LLMs takes complex data and represents it as a chain of smaller, linked components rather than one massive block of material, discarding redundant information while preserving what is relevant.

This, in turn, should result in AI that runs more quickly and consumes less energy because there is less excess material to process.

“The AI industry has accepted a false constraint for years — that powerful models require expensive infrastructure,” said Enrique Lizaso, co-founder and CEO of Multiverse Computing. “That constraint is gone. We have proven that AI can run at full performance on a smartphone, inside a sovereign data centre, on a factory floor with no cloud connection.” 

These models have been deployed across millions of devices including drones, satellites, vehicles, cameras and telecoms infrastructure, with partners and customers including Allianz, Bank of Canada, Bosch, Iberdrola, Indra, PwC and Telefónica.

CompactifAI can compress any open-source model and works on proprietary models where the developer grants access, restructuring a model’s internal weight matrices.

“The AI industry has built an extraordinary capability on top of an infrastructure model that wasn’t designed to sustain it,” said Damien Henault, Managing Director & Partner, Forgepoint Capital International, one of the round’s lead investors.

“The winners in the next phase of this multi trillion-dollar market will be the ones that make every data centre, every GPU and every enterprise deployment dramatically more efficient.”

Model compression is becoming an increasingly crowded field as AI companies grapple with the cost of running ever-larger models. Techniques such as pruning, quantization and distillation are now used in-house by major labs — Nvidia, for instance, has its own pruning and distillation pipeline, Minitron, which it has used to derive smaller Llama models using a fraction of the training resources normally required. Google, meanwhile, has developed a separate optimization technique called TurboQuant, which compresses AI systems’ memory usage at inference time rather than the model itself. 

Multiverse argues its tensor-network approach allows for deeper compression with less accuracy loss than rival techniques, and it has packaged the technology as a broader platform spanning cloud, on-premises and on-device deployment, rather than a standalone tool. 

An independent Sopra Steria evaluation of the CompatifAI’s compression method applied to Llama 3.1 8B, supported this claim, finding that it “not only significantly reduced the computational resources but also maintained the model accuracy, making the model more efficient, scalable and cost-effective”. 

The Series C round was co-led by Forgepoint Capital International, BNPP Solar Impulse Venture Fund and Bullhound Capital, with backing from HP, Orange Ventures, Scania Invest, Santander, Qatar Development Bank and several European public funds.

The funds will be used to expand the company’s library of pre-compressed AI models, continued research on algorithms, investments in sovereign AI gigafactory infrastructure and establishing a stronger presence in East Asia, Southeast Asia, the Middle East, Canada and the US.

Company info:

Reply

Avatar

or to participate

Keep Reading