
Jevstiller targets 98% agreement by learning familiar requests on your hardware while sending uncertain and audited queries upstream
Sure, Jev is cheap and quick to respond to queries, but you know what would be even better? How about a local student model that learns from your Jev requests and runs familiar ones on your hardware instead?
Jevstiller is an open source project that, as its name suggests, distills Jev’s outputs into a local model that can do some of Jev’s arguably fast work even more quickly and efficiently - no need to spend $42 per billion input tokens or wait around 300 ms for a response on requests it can handle locally, say Jevstiller’s creators.
Answering locally means no Jev tokens get spent, however cheap they may be, and answers can come in as little as 15 ms directly from a device’s CPU. The model that Jevstiller constructs on a user’s machine is small and fast, and, if its developers’ claims are to be believed, helps the system achieve 98 percent overall agreement with Jev by handling familiar requests locally and forwarding others to Jev.
Jev, as we’ve detailed in several stories recently, is a new paradigm of AI design that foregoes all the chattiness and free-form text generation that come with the typical modern LLM in favor of being able to answer just a few basic types of questions: Choice, Score, and Noul. Because it doesn’t generate a bunch of text and need to ruminate on the structure of human language, Jev is quick, cheap, and rapidly attracting developer interest as a practical application of AI for structured decision workloads where a full-blown generative LLM may be overkill.
The Jevstiller team likens it to a cache that sits in front of Jev itself, scooping up requests and determining whether they need to be routed to the actual Jev servers when the local version isn’t confident it can answer a question.
“A request is answered locally only if every question in it can be; otherwise the whole request goes to Jev with the caller’s own key, Jev’s response is returned unchanged, and each choice answer becomes a training row,” the Jevstiller team explains in a help document.
Spinning up a new Jevstiller basically means you’re setting up a model that doesn’t know anything about Jev responses, so it needs to learn how Jev responds to queries by passing all of them to Jev and gathering both questions and answers.
“Once there are enough samples per class, a background worker fits a student and a routing policy,” Jevstiller’s documentation explains. “Confident, familiar inputs are answered locally. Everything else goes to Jev.”
Say you use Jev to filter incoming emails to a particular team, or to determine priority of a particular task, for example. Once Jevstiller sees enough examples of those queries and responses, it can confidently answer the same way Jev would.
Maintaining agreement with Jev, as the Jevstiller team notes, is the key design problem they were seeking to address. They seem to think they’ve arrived at a solution in the form of regular auditing that constantly assesses Jevstiller’s responses against a random selection of queries.
“A fixed 2 percent of all requests goes to Jev regardless of what the local model thinks, with the local model’s answer recorded alongside,” the developers say, noting that such a solution is the only way to get an unbiased view of Jevstiller’s routing since it automatically declines to answer difficult questions. Jevstiller continually retrains itself based on all Jev responses, and the target agreement rate can be configured at the local level.
If the audit suggests at any time that Jevstiller’s agreement rate with Jev has fallen below the configured target, the audit rate goes up so Jevstiller’s retraining happens more quickly. If the worst happens and Jevstiller thinks its ability to agree with Jev has collapsed, "every request goes back to Jev and training restarts from that point."
The team’s own experimentation demonstrated that the concept works, as they explained.
“In the 24-hour soak, a stand-in Jev silently changed every answer at hour twelve,” Jevstiller’s devs said of an experimental run. “The local share fell from 90% to 9% within four minutes as the audits caught it, and was back at 90% within 49 minutes, trained on post-change answers only, with nobody touching anything.”
Agreement with Jev doesn’t mean a response is correct, though, which the Jevstiller team admits. “Agreement is not accuracy,” they note.
In other words, whether you go this route to even further reduce your AI bill or choose some other method, Jev is still just another AI model that can make mistakes, no matter how good it claims to be.
Those who want to give Jevstiller a try can learn more about getting started here.®