Open Inference

Contact

Berkeley, CA
Email: comms[at]openinference[dot]ai

About

Open models are optimized for GPUs. On other accelerators, they run, but not fast enough. Open Inference reshapes models to fit TPUs and Trainium: how attention is organized, how context is selected, and how weights are stored. See our models on OpenRouter.

We share the engineering as we go. To learn more or work with us, get in touch.