SESI PROGRAMMING LANGUAGE DOCS
⌕

Model, Image & Video Config in Sesi

model(), image(), and AI video() calls take an optional config block between the model name and the prompt. Config keys are unquoted identifiers.

model_call := 'model' '(' string ')' config_block? '{' prompt '}'
image_call := 'image' '(' string ')' config_block? '{' prompt '}'
video_call := 'video' '(' string ')' config_block? '{' prompt '}'
config_block := '{' key ':' value (',' key ':' value)* '}'

---

Basic Call (No Config)

let response = model("gemini-3.5-flash-lite") {"Say hello"}
show response

---

model() Config Keys

Key Type Description
thinkingLevel string Reasoning effort: "minimal", "low", "medium", "high"
max_tokens number Maximum tokens in the response
images string \ array Local file path(s) for vision input
stream bool \ fn Stream output to stdout (true) or a callback fn
response string Set to "full" to return text, finish reason, cache state, and token usage
cache bool Set to false to bypass Sesi Logic Caching
system string System instruction for Gemini, GPT, and local models
search (no value) Enable web search grounding for real-time information
temperature number ⚠️ Deprecated in Gemini 3.5+. Use thinkingLevel.
top_k number ⚠️ Deprecated in Gemini 3.5+. Use thinkingLevel.
top_p number ⚠️ Deprecated in Gemini 3.5+. Use thinkingLevel.
tools array<object> Function schemas exposed to models with tool-aware templates

---

Local Models

Use "local" to run the default quantized instruction model directly:

let answer = model("local") {max_tokens: 256, temperature: 0.3} {"Explain closures simply."}

The first call downloads the ONNX weights to ~/.cache/sesi/models. The loaded

pipeline is reused for later calls in the same process, and downloaded weights

are reused across processes.

Configure the provider with environment variables:

Variable Default
SESI_LOCAL_MODEL onnx-community/Qwen2.5-0.5B-Instruct
SESI_LOCAL_DTYPE q4
SESI_LOCAL_DEVICE cpu
SESI_LOCAL_CACHE_DIR ~/.cache/sesi/models
SESI_LOCAL_SYSTEM_PROMPT Sesi's concise local-assistant system prompt
SESI_LOCAL_WARN_TOKENS 2048

An explicit model can also be selected in the model name:

let answer = model("local:onnx-community/Qwen3-0.6B-ONNX") {"Say hello."}

Local calls support text generation, streaming, logic caching, token usage, and

tool calling for models whose chat templates support tools. Images, audio, and

search grounding are rejected with an explicit error. The system config key

overrides SESI_LOCAL_SYSTEM_PROMPT for one call.

let tools = [{
  type: "function",
  function: {
    name: "lookup_weather",
    description: "Get weather by city",
    parameters: {
      type: "object",
      properties: {city: {type: "string"}},
      required: ["city"]
    }
  }
}]

let call = model("local:onnx-community/Qwen3-0.6B-ONNX") {tools: tools} {"Use lookup_weather for New York City."}

When the model requests a tool, Sesi returns canonical JSON containing name,

args, and an optional call_id, matching hosted-model behavior.

Inputs above 2,048 tokens are allowed but emit a CPU-performance warning. See

Local Models for context limits, configuration,

and a reference benchmark.

model("local") {max_tokens: 256, temperature: 0.5, system: "Use the tone of a motivational speaker that doesn't let clients give up on their tasks."} {query}

---

thinkingLevel

Controls how much reasoning effort the model applies before responding:

// Fastest — minimal reasoning, not compatible with newer models (e.g., gemini-3.5-flash and later)
let r1 = model("gemini-3.1-pro-preview") {thinkingLevel: "minimal"} {"Summarize in one sentence:" text}

// Balanced
let r2 = model("gemini-3.8-flash") {thinkingLevel: "low"} {"Analyze this code:" code}

// Deep reasoning
let r3 = model("gemini-3.8-flash") {thinkingLevel: "medium"} {"Solve this step by step:" problem}

---

max_tokens

Cap the response length:

let brief = model("gemini-3.1-flash-lite") {max_tokens: 100} {"Explain quantum computing."}

---

images — Vision Input

Pass a local image path to give the model visual input:

// Single image
let description = model("gemini-3.5-flash-lite") {images: "photo.png"} {"Describe what you see."}

// Multiple images
let comparison = model("gemini-3.7-flash") {images: ["before.png", "after.png"]} {"What changed between these two images?"}

---

stream

Stream tokens as they arrive instead of waiting for the full response.

To stdout

let response = model("gemini-3.1-flash-lite") {stream: true} {"Write a short poem."}

// tokens show to terminal in real-time
show "Final:" response

To a callback function

fn handleChunk(chunk) {
  show "Chunk:" chunk
}

let response = model("gemini-3.1-flash-lite") {stream: handleChunk} {"Explain closures."}
show "Final:" response

The return value is always the fully accumulated response string, regardless of whether streaming is on.

---

search — Web Search Grounding

search takes no value. Adding it to the config block tells the model to ground its response in live web search results:

let response = model("gemini-3.1-flash-lite") {search} {"What is the weather in Tokyo right now?"}
show response

Combine with other keys normally:

let response = model("gemini-3.1-flash-lite") {search, max_tokens: 200} {"Latest news in Media this week."}

---

cache

Sesi caches model responses by default. Set cache: false to force a fresh call:

let fresh = model("gemini-3-flash-preview") {cache: false} {"What time is it?"}

---

Combining Config Keys

Multiple keys are comma-separated on one line:

let result = model("gemini-3.8-flash") {thinkingLevel: "medium", max_tokens: 500} {"Analyze this document:" doc}

let scan = model("gemini-3.8-flash") {images: "receipt.png", thinkingLevel: "low"} {"Extract all line items as JSON."}

---

image() Config Keys

Key Type Description
ratio string Aspect ratio — e.g. "1:1", "16:9", "4:3"
size string Output resolution — "512", "1K", "2K", "4K"
images string \ array Reference image(s) for style/context

let logo = image("gemini-3.1-flash-image") {ratio: "1:1", size: "512"} {"A minimal logo for a programming language"}
write_image("logo.png", logo)

let banner = image("gemini-2.5-flash-image") {ratio: "16:9", size: "1K"} {"A dark futuristic cityscape at night"}
write_image("banner.png", banner)

---

video() Config Keys

The AI form is video(model) {config} {prompt}. It returns Base64 MP4 data;

save it with write_file(path, value, "base64"). The local, no-AI form is

documented in VIDEO.md.

Key Aliases Type Applies to Description
ratio aspectRatio, aspect_ratio string Omni, Veo Aspect ratio: "16:9" (default) or "9:16".
images image string \ array Omni, Veo Local reference-image path or paths. Veo currently uses the first path.
task — string Omni text_to_video, image_to_video, reference_to_video, or edit. Usually omit it.
duration durationSeconds, duration_seconds number Veo Clip length: 4, 6, or 8 seconds.
resolution size string Veo "720p" (default), "1080p", or "4k". Veo Lite does not support 4K.
poll_interval pollInterval number Veo Milliseconds between generation-status checks; defaults to 10000.
negative_prompt negativePrompt string Veo Content to exclude from the generated clip.
audio generateAudio, generate_audio bool Veo Enable generated audio for model variants that support it.

let clip = video("veo-3.1-fast-generate-preview") {ratio: "9:16", duration: 8, resolution: "1080p", audio: true} {"A paper boat travels through a rain-filled neon street at night, slow tracking camera, no captions."}
write_file("generated_media/boat.mp4", clip, "base64")

For image-guided Gemini Omni Flash generation:

let clip = video("gemini-omni-flash-preview") {images: "product.png", ratio: "16:9", task: "image_to_video"} {"Slow orbit around the product on a bright studio tabletop, soft natural shadows, no text overlays."}
write_file("generated_media/product.mp4", clip, "base64")

---

Quick Reference

// No config
let r = model("gemini-3.5-flash-lite") {"Hello"}

// thinkingLevel
let r = model("gemini-3.8-flash") {thinkingLevel: "low"} {"Summarize:" text}

// max_tokens
let r = model("gemini-3.7-flash") {max_tokens: 200} {"Explain this."}

// Vision input
let r = model("gemini-3.5-flash-lite") {images: "scan.png"} {"Transcribe all text."}

// Streaming to stdout
let r = model("gemini-3.1-flash-lite") {stream: true} {"Write a poem."}

// Streaming with callback
fn onChunk(chunk) { show chunk }
let r = model("gemini-3.1-flash-lite") {stream: onChunk} {"Tell a story."}

// No cache
let r = model("gemini-3.1-flash-lite") {cache: false} {"What's trending?"}

// Combined
let r = model("gemini-3.8-flash") {thinkingLevel: "medium", max_tokens: 500, images: "doc.png"} {"Analyze this."}

// image()
let img = image("gemini-2.5-flash-image") {ratio: "16:9", size: "1K"} {"A sunset over the ocean"}
write_image("output.png", img)

// AI video()
let clip = video("veo-3.1-fast-generate-preview") {ratio: "16:9", duration: 8} {"A kite flying above a beach at sunrise."}
write_file("clip.mp4", clip, "base64")

---