Typed questions in, calibrated answers out, in milliseconds.
One binary, models pulled by name and a TypeSafe-compatible API, the way Ollama runs LLMs.
Website · Models · Results · Docs · Releases · Hugging Face
curl -fsSL https://ollaya.dev/install.sh | sh
ollaya run winnow:e4b --preset triage "I was charged twice this month and want a refund."19 model families, from millisecond encoders that run on a CPU (laya, nli, gliclass, von, decima) to
decoders that take a GPU (winnow, clef, kev, decider, nimble, jeb, jeeves, cygnet, snap, arbiter and more).
Their accuracy, calibration and speed on our own GPUs and CPUs, with the raw data, are at
ollaya.dev/results.
Weights always come from their authors' own Hugging Face repositories, pinned to a commit and verified by sha256. Ollaya never re-hosts them.
Created and maintained by Mert Cobanov (@mertcobanov). Apache-2.0. Ollaya is an independent project, not affiliated with Ollama or TypeSafe.