Sessionboxer

Sessions

Talk to the agent: attachments, dictation, Markdown and the queue

How the chat works in Sessionboxer: folded tool calls, attached files and pasted screenshots, offline dictation with whisper.cpp, the Markdown prompt box and the message queue.

Feature page: Prompts with files, a queue, dictation — screenshots, things you can do with it and how other products compare.

The chat shows the agent's messages and, folded, each tool it used: commands, file edits, and the screenshots it took while using the desktop (a folded row that holds a screenshot shows a small picture icon next to its chevron; open the row to see it). Press Enter to send, Shift+Enter for a newline. While the agent works, Send turns into Stop, which interrupts the turn; what you typed stays in the box. Hover a message for Copy (its Markdown source) and Copy rich (formatted: headings, lists, links and coloured code survive a paste into a document or mail) in its top-right corner.

While the agent works, its messages of the turn (text, thoughts, tool calls, plans) fold behind one squiggly rule — a spinner, Working… n messages so far and a live m:ss — and when the turn ends only its last message stays, the summary, with Show all n messages · 0:26 in the middle of the rule. Open it to see every message as it happened, and fold them again from either end. Every message carries its time at the bubble's bottom-right, as people say it (just now, 12 min ago, yesterday 14:32); hover it for the exact date in your browser's time zone.

The model picker at the bottom left of the prompt box switches the model for the rest of the conversation, and the pickers next to it (Claude Code: Effort, Fast mode) do the same for the agent's other settings; the agent decides which appear for the current model (Fable, for one, has no Effort or Fast mode). While the agent is working a change waits until the current turn ends (the picker shows pending), and the chat shows a Model now: … / Effort now: … marker when it takes effect. The choices survive Stop/Resume, and a setting the current model does not offer is kept for when you switch to one that does.

Attach files with the 📎 button, by dropping them on the prompt box or by pasting them (a screenshot from the clipboard works). Each file is uploaded into the box under /workspace/.sessionboxer/uploads/ before you send, with a progress bar on its chip (✕ removes it); a chip that failed says why. When you send, the prompt tells the agent the path of every file, so it can read, run or convert anything with its own tools or a shell. Images (PNG, JPEG, GIF, WebP up to 5 MB) and text files up to 64 KB are also handed to the model directly, as part of the prompt, when the agent supports that (Claude Code and Devin both do); other files, or larger ones, are reachable by path only. Up to 20 files per prompt, 512 MB each. Your message in the chat shows the files it carried, images inline and the rest as chips with a Download link. The uploads folder is left out of git status (through .git/info/exclude) and of Pull to folder, but it is part of the box like any other file (snapshots carry it); delete it in the box when you no longer need it. The New session screen has the same 📎 button, drop zone and paste: there is no box yet, so the files wait under ~/.sessionboxer/staging/ on the machine running Sessionboxer and are copied into the box as soon as it is up, before the first prompt goes out (the prompt may be just the files); files staged and then abandoned are deleted when you leave the screen, or after a day.

Dictate with the 🎤 button in the prompt box's toolbar (on the New session screen too): tap to record, tap again and the words are added to the end of your draft (nothing is sent until you press Send or Start). The clip is transcribed on the machine running Sessionboxer with whisper.cpp, offline — a phone paired through a tunnel sends its clip there too, so no account and no cloud speech service are involved. The first time, whisper-cli (a few MB, from this repository's Releases) and the model (small, 466 MB, from Hugging Face) are downloaded once into ~/.sessionboxer/; the line under the box shows the progress, then Transcribing…. Settings → Dictation picks the model — base is faster and rougher, medium (q5_0) more accurate and 2–3× slower — and the language (detect, English, Spanish, …; naming it makes transcription about twice as fast) and can download the model ahead of time or delete ones you no longer use. A 15-second prompt takes about 3 s with small on a laptop CPU. The microphone needs a secure page: localhost, or https:// through a tunnel — plain http:// over the LAN has no microphone, use the keyboard's own dictation there (Win+H, the mic key on macOS/iOS/Android), which types into the box like anywhere else.

The prompt box is Markdown: write it raw or flip the Preview switch to edit it formatted, use the toolbar for formatting either way, drag the divider to make the box taller, or go full screen with the zen button. Enqueue (Ctrl+S) puts a message in the session's Queue instead of sending it now: the queue plays by itself, each message going out as soon as the agent is idle (right away, when the current turn ends, or when a stopped session is resumed), one turn at a time. From the list you can load a message back, send one now, reorder or delete; Pause holds the queue (the current turn finishes, messages you enqueue meanwhile wait behind the others) and Resume lets it go again.

This chapter is generated from docs/GUIDE.md in the Sessionboxer repository. Found a mistake? Open an issue.