Build a Model Router with Jev, Spin & TypeScript
In this article, I will show you how to build a lightweight model router as a single WebAssembly component with Spin. Instead of hard-coding which Large Language Model (LLM) answers which request, we’ll let Jev decide, forward the prompt to the model it picked on an isolated Ollama instance, and hand the real answer back. We’ll also use Jev’s calibrated confidence score as a safety net that escalates to a stronger model whenever Jev is unsure.
What Is Jev?
Before we route anything, let’s talk about the model doing the routing.
Jev is the first System One Model from TypeSafe AI. Unlike the chat models you’re used to, Jev does not generate text. You hand it some unstructured state plus a typed question, and it hands back a typed, probabilistic decision your code can act on directly. TypeSafe frames it nicely: think of Jev as a frontier-intelligence function call - unstructured state in, typed decisions out.
That matters more than it sounds. Because the answer’s shape is defined by your schema before the call, there’s nothing to parse and no retry loop for malformed JSON. The API exposes a small set of question primitives - choice for categorical classification, score for ratings, and noul for yes/no probabilities. For a router, choice is exactly what we need: “Given this prompt, which of my models should answer it?”
Every decision also ships with a calibrated confidence value. Standard LLMs are famously overconfident, even when you ask them for a probability; Jev’s confidence instead reflects how concentrated its probability distribution is - how torn it is between your options. That single number is what turns a classifier into a router you can trust in production.
Best of all, Jev is available straight through OpenRouter as typesafe/jev-1.13. You need nothing more than an OpenRouter API key, and you’re billed per input token (output is free).
What We Will Build
Our router is a single WebAssembly component, written in TypeScript. It exposes one HTTP endpoint - POST /ask - that takes a user prompt, asks Jev which model should handle it, calls that model on Ollama, and returns the real answer.
The Mode router is based on four building blocks: Spin variables let us swap the candidate models, the fallback, the Ollama endpoint, and the API keys without re-compiling; the Jev decision call is where all the magic happens; a confidence-based safety net keeps a low-confidence guess from silently landing in front of a user; and the Ollama call forwards the prompt to the chosen model and returns a real answer.
Let me walk you through the corresponding code sections.
Configuring the Model Router
[variables]
open_router_api_key = { required = true, secret = true }
open_router_endpoint = { default = "https://openrouter.ai" }
ollama_endpoint = { default = "https://my-ollama:8080" }
ollama_api_key = { required = true, secret = true }
models = { default = "gpt-4:latest,llama3.2:1b,qwen2.5-coder:7b,deepseek-v3.2:cloud" }
fallback_model = { default = "gpt-4:latest" }
[component.jev-model-router-ts]
allowed_outbound_hosts = ["{{ open_router_endpoint }}", "{{ ollama_endpoint }}"]
[component.jev-model-router-ts.variables]
open_router_api_key = "{{ open_router_api_key }}"
ollama_endpoint = "{{ ollama_endpoint }}"
ollama_api_key = "{{ ollama_api_key }}"
models = "{{ models }}"
fallback_model = "{{ fallback_model }}"
The models variable is just a comma-separated string. On the TypeScript side, a small loadConfig function reads each variable, splits the list, and throws early if a required value is missing - so a misconfigured deployment fails loudly, not halfway through a request.
Note: For this demo, the names in models are hard-coded, and every one must be served by your Ollama instance for the forwarding step to work. In production you wouldn’t pin them like this - you’d resolve the list at compile time, or discover it in-flight via Ollama’s list-models endpoint (GET /api/tags) and feed whatever is available straight into Jev’s criteria, so the router can never pick a model Ollama can’t serve.
Making the Decision
This is the heart of the router, and it’s shorter than you might expect. Using the OpenRouter SDK, we turn our list of candidate models into Jev’s criteria, pose a single choice question, and pass the user’s prompt as state:
const res = await client.alpha.decisions.create({
decisionsRequest: {
model: 'typesafe/jev-1.13',
questions: {
model_route: {
type: 'choice',
criteria: convertModelNamesToCriteria(cfg.models),
instructions:
'Which of the following models should be used to answer the users prompt?',
},
},
state: {
user_prompt: prompt,
},
},
});
That’s it. No chat history, no system prompt gymnastics, no JSON schema wrestling. We asked a typed question and we’ll get a typed answer back - answer.choice will always be one of the models we provided, because those are the only values Jev is allowed to return.
Building a Safety Net With Confidence
Here’s where Jev earns its keep. Every answer carries a confidence value, and that’s our signal for building a cascade: trust Jev when it’s sure, and escalate to a more capable model when it isn’t.
let choice = answer.choice;
if (answer.confidence !== undefined) {
console.log(`Confidence: ${answer.confidence.toFixed(2)}`);
// Fall back to a more capable model if Jev is unsure
if (answer.confidence < 0.6) {
console.log("Low confidence detected. Escalating...");
choice = 'opus';
}
}
Pick the threshold that matches how expensive your mistakes are - a cheap internal tool can tolerate a low bar, while anything user-facing probably wants to escalate much more eagerly. And if Jev can’t produce a decision at all, we quietly fall back to the configured fallback_model rather than failing the request.
Note: The confidence number is a measure of how concentrated Jev’s probabilities are, not a guarantee of correctness. Treat it as a ranking signal and tune your threshold against your own traffic - calibration holds in aggregate, not on any single call.
Forwarding the Prompt to Ollama
With a model chosen, the last step is to actually answer the user. A small src/ollama.ts module POSTs the prompt to Ollama’s chat endpoint using the model Jev picked, and returns the generated text:
const res = await fetch(`${cfg.ollamaEndpoint}/api/chat`, {
method: 'POST',
headers: {
'content-type': 'application/json',
authorization: `Bearer ${cfg.ollamaApiKey}`,
},
body: JSON.stringify({
model,
messages: [{ role: 'user', content: prompt }],
stream: false,
}),
});
const data = await res.json();
return data.message.content;
That’s the whole forwarding path: one request with the chosen model, the prompt as a single message, and stream: false so the complete answer comes back in one shot. We pull the text from data.message.content and return it.
Exposing It Over HTTP
The HTTP surface is deliberately thin. We use the Hono router with Spin’s service-worker adapter, read the incoming prompt, call into our decision logic, forward the prompt to the chosen model on Ollama, and return the generated answer:
app.post('/ask', async (c: Context) => {
const cfg = loadConfig();
const req = await c.req.json<RequestModel>();
const decision = await decideWhichModelToUse(req.prompt, cfg);
const answer = await askOllama(req.prompt, decision.model, cfg);
return c.json({ model: decision.model, answer });
});
Validation, error handling, and logging round it out, but the shape stays this simple. That’s the point - the router is a thin, well-behaved HTTP component, Jev does the routing, and Ollama produces the answer.
Running It Locally
Testing a Spin app on your local machine is a single command. You compile to WebAssembly and start the component while supplying your configuration, all in one go:
spin up --build \
--variable open_router_api_key="{your_openrouter_api_key}" \
--variable ollama_endpoint="http://my-ollama:8080" \
--variable ollama_api_key="{your_ollama_api_key}"
Make sure your Ollama instance is running and already serving every model listed in models - the router can only forward to a model Ollama actually has.
Once Spin reports that your component is listening on port 3000, send it a prompt from a second terminal:
curl -iX POST \
-d '{ "prompt": "Write a Python function to reverse a linked list" }' \
-H 'content-type:application/json' \
http://localhost:3000/ask
This time you get back a JSON payload telling you both which model Jev routed to and the answer it generated:
{
"model": "qwen2.5-coder:7b",
"answer": "def reverse_linked_list(head):\n prev = None\n ..."
}
With the \n escape chars rendered, the answer reads like this:
def reverse_linked_list(head):
prev = None
while head:
head.next, prev, head = prev, head, head.next
return prev
The logs show Jev’s chosen route and its confidence for every request. Send a coding prompt and a casual chat prompt back to back, and you’ll see each one routed to a different model - exactly the behaviour we were after.
Deploying to Akamai Functions
Once the router works locally, going global is a single command. Authenticate first with spin aka login, then deploy:
spin aka deploy --build \
--create-name jev-model-router \
--variable open_router_api_key="{your_openrouter_api_key}" \
--variable ollama_endpoint="https://my-ollama:8080" \
--variable ollama_api_key="{your_ollama_api_key}"
The --build flag compiles the component to WebAssembly as part of the deploy, and we pass the same variables we used locally. When it finishes, the command prints the public endpoint where your router now lives - globally distributed and fully managed. Point the earlier curl at that endpoint instead of localhost to try it.
Get the Code
I’ve kept this to the highlights on purpose - the full application is ready for you to clone and run:
Open sourceReady to RunGrab the full sample code and run it yourself!akamai-developers/jev-model-routerGrab an OpenRouter API key, point it at a running Ollama instance, and you’ll have a working model router in under a minute.
Recap
We built a model router that:
- runs as a single WebAssembly component with Spin
- hands the routing decision to Jev
- uses Jev’s calibrated confidence to escalate low-confidence prompts to a stronger model
- forwards the prompt to the chosen model on Ollama
System One Models like Jev are a genuinely new building block: fast, typed, calibrated decisions software can act on without parsing a token of free-form text. Pairing that with the portability of a WebAssembly component feels like a very natural fit.
Are you routing between models, or using Jev for something else entirely? Join the Edge Case, our developer community over on Discord, and let me know - I’m curious what you’re building.