Local Models
Sesi runs local text models directly in Node.js:
let answer = model("local") {max_tokens: 256} {"Explain closures."}
A provider API key is not required. The default model is onnx-community/Qwen2.5-0.5B-Instruct with the q4 CPU
backend. Model weights are downloaded once and cached locally.
Context and performance thresholds
The default model has a 32,768-token context window shared by the system
prompt, user input, chat-template overhead, and generated output.
Sesi's recommended CPU threshold is 2,048 input tokens. Calls above
that threshold are allowed, but Sesi prints a warning because latency and
memory use increase sharply with longer prompts.
let tokens = count_tokens(document, "local")
if tokens > 2048 {
show "Consider chunking this document before local inference."
}
Configure or disable the warning:
# Warn above 4,096 input tokens
export SESI_LOCAL_WARN_TOKENS=4096
# Disable the warning
export SESI_LOCAL_WARN_TOKENS=0
The exported runtime constants are:
DEFAULT_LOCAL_MODELDEFAULT_LOCAL_MODEL_WARNING_TOKENS
Reference benchmark
Observed on July 28, 2026:
| Item | Value | |
|---|---|---|
| Computer | MacBook Air (Apple M2) | |
| CPU | 8 cores: 4 performance, 4 efficiency | |
| Memory | 8 GB | |
| Architecture | arm64 | |
| Operating system | macOS 26.6, build 25G5043d | |
| Node.js | 24.12.0 | |
| Runtime | Transformers.js 4.2.0, ONNX q4 CPU | |
| Input | README.md, 18,829 characters / 4,612 tokens | |
| Output cap | 256 tokens | |
| Result | No response returned within five minutes; call stopped manually |
This is a single reference point, not a universal benchmark. Performance varies
with hardware, model, dtype, device, prompt length, and output length. On the
reference 8 GB M2 system, keep interactive prompts below 2,048 tokens and chunk
larger documents.
Configuration
| Variable | Default | |
|---|---|---|
SESI_LOCAL_MODEL |
onnx-community/Qwen2.5-0.5B-Instruct |
|
SESI_LOCAL_DTYPE |
q4 |
|
SESI_LOCAL_DEVICE |
cpu |
|
SESI_LOCAL_CACHE_DIR |
~/.cache/sesi/models |
|
SESI_LOCAL_SYSTEM_PROMPT |
Sesi's local-assistant prompt | |
SESI_LOCAL_WARN_TOKENS |
2048 |
Use model("local:organization/model") to select a different compatible ONNX
model for one call.
Tool calling
Local models can receive function schemas through the tools model config when
their tokenizer chat template supports tool use:
let tools = [{
type: "function",
function: {
name: "lookup_weather",
description: "Get weather by city",
parameters: {
type: "object",
properties: {city: {type: "string"}},
required: ["city"]
}
}
}]
let call = model("local:onnx-community/Qwen3-0.6B-ONNX") {tools: tools} {"Use lookup_weather for New York City."}
Tool-aware local output is normalized to the same JSON shape as hosted
providers:
{"name":"lookup_weather","args":{"city":"New York City"}}
Sesi supports structured tool-call messages and the tagged <tool_call>...</tool_call>
format used by Qwen chat templates. Models without a tool-aware chat template
may ignore the supplied schemas and return ordinary text.