Speech Tools
Sesi can speak and transcribe audio files.
Speak text
speech() sends text to your operating system's installed voice engine.
speech("Hello from Sesi")
speech("Bonjour depuis Sesi", "Thomas")
The optional voice name is platform-specific. Sesi uses say on macOS, System.Speech on Windows, and espeak-ng on Linux.
Pass a Gemini TTS model as the third argument to return base64 audio instead:
let audio = speech("Hello from Sesi", "Kore", "gemini-2.5-flash-preview-tts")
// Save the returned base64 WAV audio to disk
write_file("speech.wav", audio, "base64")
Transcribe audio
from_speech() transcribes an audio file using nodejs-whisper.
let transcript = from_speech("standup.mp3", "en")
show transcript
A downloaded model must be installed (npx nodejs-whisper download base.en).
Pass a Gemini model as the third argument to use it instead of Whisper.