Hyper Efficiency
Sixteen bits of precision cut to one, with the model still doing its job.
We research 1-bit AI models
By default, most AI models are trained in full precision. This means each parameter within the model has 16 bits. This works, but we ask the question: does every parameter really need 16 bits?
We research 1-bit AI models, where each parameter is either a 0 or 1*. We also study tenary quantization and other model compression and quantisation techniques.
Sixteen bits of precision cut to one, with the model still doing its job.
A model that needed 32GB fits in 2GB — small enough for a phone.
Multiplies become additions, so each token takes far less arithmetic.
Less memory traffic, less energy per token, no data centre required.
Need a 1-bit model?
GET IN TOUCH ↗