Reference
Every tool of the built-in desktop and sessionboxer MCP servers
The two MCP servers every session's agent gets: desktop (screenshot, zoom, mouse, keyboard, scroll, wait, recordings with captions and narration) and sessionboxer (whoami, docs, PRs, snapshots, queue, Auto QA, notifications, panes, other sessions, schedules), each tool with its parameters, limits and what it returns.
Feature page: The agent knows where it is — screenshots, things you can do with it and how other products compare.
Every session's agent gets two MCP servers from Sessionboxer itself, next to the ones you add in Global settings: desktop, which sees and drives the box's screen, and sessionboxer, which tells the agent which session it is and lets it act on Sessionboxer. Both are always on (Advanced… → MCP & connectors shows them as switches you cannot turn off; only the sessionboxer policy changes). This page lists each tool as the agent sees it: the exact name, every parameter with its accepted values and defaults, what comes back, and what happens on the Sessionboxer side. The sources are packages/computer-use-mcp and packages/sessionboxer-mcp in the repository; the user guide says how the same things look from the UI.
Conventions: parameters marked required have no default. Sizes are in characters. A tool that fails (a parameter out of range, a policy that forbids it, a session that does not exist) returns an MCP tool error whose text says why, so the agent can tell you rather than guess.
Where the servers run
Both servers are stdio MCP processes inside the Linux box of the session (the Sandbox container). The agent's MCP configuration (.mcp.json for Claude Code, the equivalent for Codex, Cursor and Devin) lists them under the names desktop and sessionboxer.
| Environment | Agent, repositories, your MCP servers | desktop |
sessionboxer |
|---|---|---|---|
| Docker · Linux | In the box (/workspace) |
The box's own X display, 1024×768 by default (SESSIONBOXER_DISPLAY_WIDTH / _HEIGHT) |
In the box; talks to the box's daemon |
| QEMU · Windows | In the Windows VM (C:\workspace), agent started over SSH |
Still in the Linux box: the display shows the VM's desktop full screen over RDP, and the tools act on that picture. The agent in the VM reaches the server through a bridge command the daemon installs | Same bridge; the daemon and the Control Plane connection are on the Linux side |
| QEMU · macOS | In the macOS VM (/Users/agent/workspace), agent started over SSH |
As Windows, over VNC | Same |
What this means for a Windows or macOS session: screenshots are of the remote-desktop stream, full screen at the display's size (the guest runs at 1024×768 too), typing is paced at 40 ms per key (SESSIONBOXER_TYPE_DELAY_MS, 12 ms on Linux) because RDP and VNC drop keys at xdotool's default pace, and a few combinations may be kept by the viewer instead of reaching the guest (Win-key shortcuts over RDP; on macOS ⌘ is super and most combinations pass, Spotlight is the fallback). The agent's briefing says so and tells it to open programs from the Start menu or Spotlight. Recordings are of the same display, so they work unchanged; their paths are Linux paths under /workspace, which the daemon mirrors into the guest's Workspace so the file shows up for the agent and in the chat.
The desktop server
The desktop MCP mirrors Anthropic's computer-use tool set, implemented with xdotool and ffmpeg on the box's X display. Coordinates are integer pixels, [x, y], origin at the top-left of the screen; the screenshot tool's description tells the agent the screen size. Every tool that moves the pointer to a target glides there (an eased motion of 100–300 ms, about 120 positions a second) so hover effects and drag-and-drop see a moving pointer; SESSIONBOXER_MOUSE_GLIDE=0 in the server's environment makes it jump.
Tools that act (move, click, type, scroll…) return the text OK; screenshot, zoom and wait return a PNG image; the recording tools return JSON.
| Tool | Does |
|---|---|
screenshot |
Picture of the whole screen |
zoom |
Picture of a region, scaled up |
cursor_position |
Where the pointer is |
mouse_move |
Move the pointer |
left_click, right_click, middle_click, double_click, triple_click |
Click |
left_click_drag |
Drag from one point to another |
left_mouse_down, left_mouse_up |
Press and release separately |
type |
Type text |
key |
Press keys and combinations |
hold_key |
Hold a key for a while |
scroll |
Wheel scrolling |
wait |
Pause, then a screenshot |
start_recording |
Begin an .mp4 recording of the screen |
annotate_recording |
Caption the running recording |
stop_recording |
Finish the video: condense, captions, narration |
narrate_recording |
Add spoken narration to a finished video |
recording_status |
Is a recording running |
Seeing the screen
screenshot
Take a screenshot of the whole desktop. The agent is told to call it before acting and after any action whose result it needs to see, since the other tools only answer OK.
No parameters. Returns a PNG of the full screen.
zoom
Capture a rectangular region and scale it up to the full screen size, to read small text or inspect a detail. Coordinates the agent sees in the zoomed image are not screen coordinates; it maps them back through the region.
| Parameter | Type | Meaning |
|---|---|---|
region |
[x0, y0, x1, y1], integers, required |
Top-left and bottom-right corners, in screen coordinates |
Returns a PNG.
cursor_position
No parameters. Returns { "x": …, "y": … }, the pointer's current position.
Mouse
mouse_move
Move the pointer to a coordinate without clicking.
| Parameter | Type | Meaning |
|---|---|---|
coordinate |
[x, y], integers, required |
Where to go |
left_click, right_click, middle_click, double_click, triple_click
Click the left, right or middle button; double-click; triple-click (selects a line or paragraph in most programs). All five take the same parameter.
| Parameter | Type | Meaning |
|---|---|---|
coordinate |
[x, y], integers, optional |
Where to click; omitted = click where the pointer already is |
left_click_drag
Press the left button at start_coordinate, move to coordinate, release. The move between the two is a glide, so drop targets that watch the pointer travel see it.
| Parameter | Type | Meaning |
|---|---|---|
start_coordinate |
[x, y], integers, required |
Where the button goes down |
coordinate |
[x, y], integers, required |
Where it is released |
left_mouse_down, left_mouse_up
Press and hold the left button, and release it, as two calls, for drags that need something in between (a mouse_move through several points, a key while dragging). Both take the optional coordinate of the click tools: the pointer moves there first.
Keyboard
type
Type a string at the current focus, as a keyboard would. For shortcuts and special keys the agent uses key.
| Parameter | Type | Meaning |
|---|---|---|
text |
string, at least 1 character, required | What to type |
Typed in chunks of 50 characters with a pause between keys (12 ms on a Linux desktop, 40 ms on a Windows or macOS one).
key
Press a key or a combination, named the way xdotool names them: Return, Escape, Tab, BackSpace, Delete, Home, End, Page_Down, Up, F5, ctrl+s, alt+Tab, super, ctrl+shift+t. Several presses in a row are separated by spaces ("ctrl+a Delete").
| Parameter | Type | Meaning |
|---|---|---|
text |
string, at least 1 character, required | The key or combination(s) |
hold_key
Hold a key or combination down for a number of seconds, then release it (for a program that reacts to a long press, or to keep a modifier down during another action).
| Parameter | Type | Meaning |
|---|---|---|
text |
string, required | The key or combination, xdotool names |
duration |
number > 0, at most 30, required | Seconds to hold |
Scrolling and waiting
scroll
Scroll the mouse wheel at a coordinate.
| Parameter | Type | Meaning |
|---|---|---|
coordinate |
[x, y], integers, optional |
Where to scroll; omitted = where the pointer is |
scroll_direction |
up · down · left · right, required |
Which way |
scroll_amount |
integer 1–50, default 3 | Wheel clicks |
wait
Wait for a page to load or an animation to finish, then take a screenshot.
| Parameter | Type | Meaning |
|---|---|---|
duration |
number > 0, at most 60, default 2 | Seconds to wait |
Returns a PNG of the screen after the wait.
Recording
A recording is ffmpeg grabbing the display into an H.264 .mp4 (yuv420p, faststart, so browsers play it). It runs detached from the MCP process and its state lives in a file on the box's tmpfs, so one recording runs at a time per session and a stop_recording from a later turn still finds it. The lifecycle the agent follows: start_recording → (annotate_recording, then the step, repeated) → stop_recording → mention the returned path in the reply, and the chat shows a player with the captions as clickable steps. Auto QA runs use the same tools and hand the path to e2e_finish.
start_recording
Start recording the desktop until stop_recording. Fails when a recording is already running.
| Parameter | Type | Meaning |
|---|---|---|
path |
string, optional | Output file, under /workspace and ending in .mp4; default recordings/<timestamp>.mp4 (relative paths resolve against /workspace). Anything else is refused |
fps |
integer 1–30, default 15 | Frames per second |
Returns { path, startedAt, seconds: 0, captions: 0 }.
annotate_recording
Add a caption to the running recording at this moment: one short sentence saying what the agent is about to do or what the screen now shows ("Submitting the form with an empty email"). Each caption stays on screen until the next one. The agent calls it right before each step.
| Parameter | Type | Meaning |
|---|---|---|
text |
string, 1–300 characters, required | The caption |
Returns the caption with its time from the start of the recording ({ at, text, path }). Fails when nothing is recording.
stop_recording
Stop the recording and finish the video. In order: the captions are burned into a band added under the desktop (nothing on screen is covered; the band fits about three lines and grows upwards for a longer caption); static stretches are condensed; the .vtt is written with the re-timed captions; narration is added when it applies. The tool result reports the recorded and the final length.
| Parameter | Type | Meaning |
|---|---|---|
condense |
boolean, default true |
Cut every stretch where nothing changes on screen (a page loading, a build, the agent thinking) down to a hold of hold_seconds, so waiting does not pad the video while each state stays on screen long enough to read. Motion plays at real speed. false keeps the real timing, for animations or performance demos |
hold_seconds |
number 0.5–10, default 1.5 | How long a static stretch stays after condensing |
captions |
both · burn · track · none, default both |
Burn the captions into the frames, write them as a WebVTT file next to the video (demo.vtt beside demo.mp4, which the chat's player offers as a subtitle track), both, or drop them |
narrate |
boolean, optional | Force narration on or off for this video (because you just asked for it, say); omitted = follow Global settings → MCP & connectors → desktop → Narrate recordings |
narration_language |
en · en-gb · es · fr · hi · it · pt, optional (default en) |
The language the captions are written in; picks the default voice for it unless narration_voice is given |
narration_voice |
a Kokoro voice name, optional | The first letter is the language (a en-US, b en-GB, e es, f fr, h hi, i it, p pt-BR), the second f/m. Accepted: af_alloy af_aoede af_bella af_heart af_jessica af_kore af_nicole af_nova af_river af_sarah af_sky am_adam am_echo am_eric am_fenrir am_liam am_michael am_onyx am_puck am_santa bf_alice bf_emma bf_isabella bf_lily bm_daniel bm_fable bm_george bm_lewis ef_dora em_alex em_santa ff_siwis hf_alpha hf_beta hm_omega hm_psi if_sara im_nicola pf_dora pm_alex pm_santa. Defaults per language: af_heart, bf_emma, ef_dora, ff_siwis, hf_alpha, if_sara, pf_dora |
narration_speed |
number 0.7–1.5, default 1 | Speaking rate |
Returns JSON:
{
path, startedAt,
recordedSeconds, // wall-clock length of the recording
seconds, // length of the finished video
condensed, // whether static stretches were cut
bytes,
captions: [{ at, text }], // with their times in the finished video
track, // the .vtt path, when one was written
narration, // see below; absent when there were no captions
warning // what did not happen as asked; the video is still usable
}
narration is one of:
{ added: true, language, voice, speechSeconds, processingSeconds }: the captions were spoken (Kokoro, a text-to-speech model inside the box; nothing leaves the machine) and muxed in as an audio track. A step shorter than its sentence holds its last frame until the sentence ends, so the video may grow a little;captionsand the.vttcarry the new times.{ pending: true, estimatedSeconds, language, voice, nextStep }: the setting is Ask when it takes longer than N seconds (the default, N = 5) and this video is above it. The silent video is delivered and the agent is told to ask you first, then callnarrate_recordingif you want it.{ skipped: "…" }: the setting is Never,narrate: falsewas passed, or the TTS model is not in the image.
Fails when nothing is recording.
narrate_recording
Add spoken narration to a finished recording: after a pending answer and your yes, or when you ask for narration later. Rewrites the .mp4 in place with the audio track (extending short steps as above) and returns the same shape as stop_recording minus the recording-time fields, with the re-timed captions and .vtt.
| Parameter | Type | Meaning |
|---|---|---|
path |
string, required | The recording's path, as stop_recording returned it |
narration_language, narration_voice, narration_speed |
as in stop_recording |
recording_status
No parameters. Returns { path, startedAt, seconds, captions } for the running recording, or { "recording": false }.
The sessionboxer server
How a call travels, and what the agent knows without asking
The box has no route to Sessionboxer's API and no token. The sessionboxer MCP posts each call to the box's own daemon (POST http://127.0.0.1:7000/sessionboxer; the Auto QA tools use /e2e), the daemon forwards it over the WebSocket the Control Plane already holds to the box, and the Control Plane, which knows from that socket which session is talking, applies the session's policy, performs the action and answers. A box can only ever speak for its own session; the agent cannot claim to be another one. When the box's daemon is unreachable the tool error says so.
Before any tool call the agent already knows where it is: its briefing opens with the session's name and environment, and .sessionboxer/session.json in the Workspace (/workspace, C:\workspace or /Users/agent/workspace) holds the id, title, URL, provider, model, environment, guest OS, creation time, what it was forked from, which session's agent created it, the conversation branch, the snapshot count, the Sessionboxer version and the policy in force, kept current by the box.
Results are JSON text. Every action that changes something (a PR attached, a snapshot, a rename, a queued prompt, a verification run, a notification, a pane opened, a session created, forked, messaged or stopped, a schedule) is also written into the transcript as a small marker at the point where it happened ("attached PR #12 (owner/repo)", "took Snapshot 3", "sent a message to Backend tests"), so the chat says what the agent did.
Policy, approval and limits
Two settings decide what the agent may do: the default in Global settings → MCP & connectors → sessionboxer and the session's own value in Advanced… / Session settings → MCP & connectors. A change applies to the running agent at its next turn.
| Policy | The agent gets |
|---|---|
| Off | No sessionboxer server at all (the briefing and session.json remain). Every call errors with "the sessionboxer tools are off for this Session" |
| This Session only | The self-knowledge tools, the tools on its own session (PRs, snapshot, queue, title, verify, notify, terminals, ui_open), session_fork of itself, approval_wait, schedules that target itself, and the Auto QA tools. A cross-session tool errors with the reason and where you can allow it |
| All Sessions (the default for new installs) | Everything, including sessions_list, session_get, session_create, session_message, session_wait, session_stop and schedules that prompt other sessions |
When the Agent creates a Session (same place): Ask me (default) or Do not ask. With Ask me, session_create returns { pending: true, id, summary, expiresAt, hint } and a card appears in your chat ("Claude Code wants to create a Session 'Backend tests'") with Allow and Deny; the agent waits with approval_wait. A card nobody answers in 10 minutes expires and counts as denied. Settled cards stay in the transcript with a link to the session they created.
Limits that hold whatever the policy: at most 3 alive sessions created by one agent at a time (stopped ones do not count; session_stop frees a place) and a global cap on agent-created sessions (Global settings → MCP & connectors → sessionboxer, 10 by default); a chain of agents prompting agents stops after 4 hops; one message in flight per sender and target (wait for the reply first); a session cannot message itself; session_stop only stops sessions the calling agent created. Sessions an agent created show child of … under their name in the sidebar and in sessions_list.
The wait tools (session_wait, approval_wait) block for at most 20 s per call (default 15) because the bridge itself times out at 20 s; the agent calls them again while the answer is still pending.
| Tool | Does | Policy |
|---|---|---|
whoami |
Everything about this session | session |
docs |
Look a topic up in the user guide | session |
settings_get |
The settings in force, secrets removed | session |
pr_attach |
Attach a pull request to the session | session |
pr_list |
The attached PRs | session |
pr_items |
A PR's review comments and checks | session |
pr_mark_addressed |
Mark items as dealt with | session |
snapshot |
Snapshot the box now | session |
queue_add |
Queue a prompt for itself | session |
queue_list |
The queue | session |
title_set |
Rename the session | session |
verify |
Open an Auto QA run with a brief | session |
notify |
Push a notification to you | session |
terminal_list |
Your Terminals | session |
terminal_read |
What a Terminal printed | session |
ui_open |
Open a pane in your browser | session |
sessions_list |
Every session | all |
session_get |
One session and its last reply | all |
session_create |
Start a new session | all |
session_fork |
Fork this session, optionally with a handoff | session |
session_message |
Prompt another session's agent | all |
session_wait |
Wait for another session's turn | all |
session_stop |
Stop a session it created | all |
approval_wait |
Wait for your Allow / Deny | session |
schedule_create |
Create a scheduled task | session (own), all (others) |
schedule_list |
The scheduled tasks | session |
e2e_plan |
Register the cases of the Auto QA run | session |
e2e_case_start |
Start a case | session |
e2e_case_end |
Record a case's result | session |
e2e_finish |
Close the run with the video | session |
Self-knowledge
whoami
No parameters. Returns the session as the Control Plane sees it, live:
{
id, title, url, provider, model, environment, // "docker-linux" | "qemu-windows" | "qemu-macos"
guest, // { os: "windows" | "macos", workspace } or null
createdAt,
forkedFrom, // { sessionId, title, snapshotId } or null
createdBy, // { sessionId, title } or null: the agent that created this session
branch, snapshotCount, sessionboxerVersion,
agentTools, // "off" | "session" | "all"
status, // running, idle, stopped…
usage, // the provider's usage windows, when it reports them
context, // context window use of the current conversation
queueLength,
panes, // which panes you have open right now
terminals: [{ id, createdAt, exitCode }],
prs: [ summaries of the attached pull requests ],
verification, // { id, status, brief } of the current Auto QA run, or null
repos: [{ name, path }] // /workspace/<name> as the agent sees it
}
docs
Look a topic up in the Sessionboxer user guide, which ships in the box (/opt/sessionboxer/docs/GUIDE.md). The agent's briefing tells it to answer questions about Sessionboxer from this rather than from memory.
| Parameter | Type | Meaning |
|---|---|---|
query |
string, 1–200 characters, required | A heading, a feature name or a few keywords |
Returns { query, sections: [{ heading, path, text }] }, the three best-matching sections by word overlap with the heading and text, each cut at 6,000 characters, path being the chain of headings ("Using it › MCP servers"). With no match, sections is empty and headings lists every heading of the guide so the agent can retry with one.
settings_get
No parameters. Returns the public settings that apply to this session (the same object the UI reads: environments available, defaults, git identity, the MCP registry without its secrets, the sessionboxer policy…), with any secret-shaped value removed before it leaves the Control Plane.
This session
pr_attach
Attach a pull request to the session: you see it in the PRs pane with its checks and review comments from then on, and the session gets notified when reviewers comment or a check fails. The agent is told to call it right after creating a PR.
| Parameter | Type | Meaning |
|---|---|---|
ref |
string, 1–500 characters, required | The PR's URL, or owner/repo#number (GitHub, or a Bitbucket Data Center URL) |
Returns the attached PR's summary. Marker in the transcript: "attached PR #n (owner/repo)".
pr_list
No parameters. Returns the pull requests attached to the session with their state, check counts and unseen review items.
pr_items
| Parameter | Type | Meaning |
|---|---|---|
pr |
string, required | The PR's id from pr_list, or its URL / owner/repo#number |
Returns the review comments, check results and other items of the PR, each with its id and whether it was marked addressed.
pr_mark_addressed
| Parameter | Type | Meaning |
|---|---|---|
pr |
string, required | As above |
items |
array of 1–200 item ids from pr_items, required |
What was dealt with |
The items show as addressed in the PRs pane.
snapshot
No parameters. Takes a Snapshot of the Sandbox now (its disk, the Workspace and the conversation), the same as the Snapshot button; you can fork from it or revert to it. Returns { id, ordinal, sizeBytes, createdAt }. Marker: "took Snapshot n". Not available in Windows and macOS sessions (their VM disk is outside the Sandbox image).
queue_add
Queue a prompt for the agent itself: it is sent as the next user turn once the current one ends, after anything you already queued. Meant for a follow-up the agent wants a fresh turn for; you see it in the queue and can edit or drop it before it goes.
| Parameter | Type | Meaning |
|---|---|---|
text |
string, 1–20,000 characters, required | The prompt |
Returns { id, position }. Marker: "queued a message (…)".
queue_list
No parameters. Returns the prompts queued for the session, in order, with their ids.
title_set
| Parameter | Type | Meaning |
|---|---|---|
title |
string, 1–200 characters, required | The new name (sidebar and browser tab) |
Marker: "renamed the Session to …".
verify
Open an Auto QA (verification) run for the work of this turn, with a brief of what it checks that the Auto QA pane shows above the cases; then the agent follows the e2e-verification skill (e2e_plan unless the cases were given here, e2e_case_start / e2e_case_end, e2e_finish). The agent is told not to call it when Sessionboxer already asked it to verify the turn (the Auto QA setting does that after every user turn).
| Parameter | Type | Meaning |
|---|---|---|
brief |
string, 1–2,000 characters, required | What the run verifies and how, in two or three sentences |
cases |
array of at most 10 { title (1–200), steps (≤4,000), expected (≤2,000) }, optional |
The cases, when already known; otherwise e2e_plan comes next |
Returns { id, status, cases: [{ index, title, status }] }. Marker: "opened verification run …".
notify
Send you a short notification about the session: a browser push (when the device is enrolled under Devices) and the bell in Sessionboxer. For something that cannot wait for the reply, not for progress.
| Parameter | Type | Meaning |
|---|---|---|
text |
string, 1–500 characters, required | The notification |
Returns { ok: true }. Marker: "notified you: …".
terminal_list
No parameters. Returns the Terminals of the session, yours and the ones ui_open opened for the agent, with id, createdAt and exitCode (null while the shell still runs).
terminal_read
| Parameter | Type | Meaning |
|---|---|---|
id |
string, required | A Terminal id from terminal_list |
lines |
integer 1–2,000, default 100 | How many of the last lines |
Returns { id, text, exitCode }: the last lines of the Terminal's retained output (what you would see scrolling up in the pane) and the shell's exit code, null while it runs.
ui_open
Ask your browser to show a pane of this session. Honoured only when you are looking at this session and not typing; otherwise it is ignored, so an agent cannot pull you away from another session or interrupt a message. With pane: "terminal" and a terminal.command, a new Terminal opens and runs the command where you can watch it (a dev server, a test run).
| Parameter | Type | Meaning |
|---|---|---|
pane |
chat · desktop · code · terminal · context · prs · e2e · schedules, required |
Which pane (e2e is Auto QA) |
terminal |
{ command: string (1–4,000) }, optional |
With terminal: the command the new Terminal runs |
Returns { pane, terminalId, shown }, shown being whether a browser of yours had the session open to receive it; terminalId can be read back with terminal_read. Marker: "opened the … pane" / "opened a Terminal running …".
Other sessions (policy All Sessions)
Session ids in these tools are the ids sessions_list returns; a prefix of 6 or more characters is enough when it is unambiguous.
sessions_list
No parameters. Returns every session on this Control Plane: id, title, url, status, provider, environment, repos (names), createdAt, createdBy, forkedFrom, queueLength, with the calling session marked self: true and the ones its agent created mine: true.
session_get
| Parameter | Type | Meaning |
|---|---|---|
id |
string, 1–100 characters, required | The session |
Returns that session's summary plus the last thing its agent said (capped in length). Never another session's whole conversation.
session_create
Start a new session whose first prompt is first_prompt; it is marked as created by this session (child of …). The new agent shares nothing with the caller but that text, so the prompt has to be the whole task, self-contained.
| Parameter | Type | Meaning |
|---|---|---|
title |
string, 1–200 characters, optional | Defaults to the start of first_prompt |
provider |
claude-code · devin · codex · cursor, optional |
Defaults to the caller's provider; the provider has to be connected |
repos |
array of at most 20 repositories, default [] |
Each { name?, source } with source either { type: "git", url, ref? } or { type: "copy", path } (a directory on the host, copied in). name is the directory under /workspace, derived from the source when omitted |
first_prompt |
string, 1–20,000 characters, required | What the new agent is asked first |
Returns { pending: false, sessionId, title, url }, or with approval on { pending: true, id, summary, expiresAt, hint } where id is the approval to pass to approval_wait. The new session takes the Global defaults (environment Docker · Linux, model, instructions, policy). Errors: 3 alive children already, the global cap reached, provider not connected. Marker: "created Session …", linking to it.
session_fork
Fork this session from a Snapshot taken now: the fork has the same files, repositories and tools. Works under the This Session only policy too, since it acts on the caller's own session.
| Parameter | Type | Meaning |
|---|---|---|
conversation |
continue · new · handoff, default continue |
continue keeps the conversation (same provider only); new starts an empty one; handoff starts from document |
provider |
claude-code · devin · codex · cursor, optional |
Another agent for the fork; then conversation must be new or handoff |
title |
string, 1–200 characters, optional | |
document |
string, 1–200,000 characters, optional | With handoff: the handoff document, written by the agent in Markdown (goal, state of the work, decisions, open items, files, how to run it). Because the agent writes it itself, the fork starts at once, without the hidden handoff turn the UI's Hand off performs |
first_prompt |
string, 1–20,000 characters, optional | A prompt queued for the fork's agent after it starts |
Returns { pending: false, sessionId, title, url, snapshotOrdinal }. Counts as a child for the 3-alive limit. Not available in Windows and macOS sessions (no snapshots there).
session_message
Send a prompt to another session's agent. Its transcript shows the prompt as from Session X (its Agent), never as your words, and the sender's shows sent a message to Y.
| Parameter | Type | Meaning |
|---|---|---|
id |
string, 1–100 characters, required | The target session |
text |
string, 1–20,000 characters, required | The prompt |
when |
now · queue, default now |
now prompts at once when the target is idle and queues it when it is busy; queue always queues it behind whatever is already there |
Refused when: the target is the caller itself (queue_add is for that), a message from this sender to that target is still in flight (session_wait for the reply first), the chain of agents prompting agents would go past hop 4, the target is in error, or the target is stopped and when is now (queue leaves it for when you resume it). Returns { sessionId, title, delivery } with delivery prompted or queued.
session_wait
Wait until another session's agent finishes its turn and its queue is empty, or the timeout passes.
| Parameter | Type | Meaning |
|---|---|---|
id |
string, 1–100 characters, required | The session |
timeout_s |
integer 1–20, default 15 | Seconds to wait at most; the call returns earlier when the session settles |
Returns { sessionId, title, still_running, status, lastReply }; the agent calls again while still_running is true. Waiting for itself is refused.
session_stop
| Parameter | Type | Meaning |
|---|---|---|
id |
string, 1–100 characters, required | A session this agent created (session_create or session_fork) |
Stops its Sandbox (you can resume it) and frees one of the 3 child places; returns { sessionId, title, status }. Any other session is yours to stop, and the tool refuses it; so is stopping itself (the agent ends its turn instead).
approval_wait
Wait for your answer to a pending approval.
| Parameter | Type | Meaning |
|---|---|---|
id |
string, 1–100 characters, required | The approval id session_create returned |
timeout_s |
integer 1–20, default 15 | Seconds to wait at most |
Returns the approval: { id, kind, summary, status, expiresAt, result?, error? } with status one of pending, allowed, denied, expired; with allowed, result is what the approved action returned (the created session). The agent is told to call again while pending and to tell you what it is waiting for; its turn may also end, in which case the card in the chat still creates the session when you allow it.
Scheduled tasks
schedule_create
Create a scheduled task, the same as one made on the Scheduled tasks page: on a cron schedule, prompt a session or start a new one each time. Prompting the caller's own session works under This Session only; another session's needs All Sessions.
| Parameter | Type | Meaning |
|---|---|---|
name |
string, 1–200 characters, required | |
cron |
string, 1–200 characters, required | A 5-field cron expression, e.g. 0 9 * * 1-5 |
timezone |
string, 1–100 characters, optional | An IANA time zone (Europe/Madrid); the Control Plane's when omitted |
action |
one of the two objects below, required |
action:
{ type: "prompt", sessionId?, text }: sendtext(1–20,000 characters) to a session at each run;sessionIddefaults to the caller.{ type: "new_session", title?, provider?, repos, prompt, stopAfter }: start a session at each run withprompt(1–20,000) as its first prompt;reposas insession_create(default[]);stopAfter(defaulttrue) stops the session once its first turn ends.
Returns { id, name, cron, timezone, nextRunAt }. The schedule is enabled at once and skips runs missed while Sessionboxer was down (the skip policy; you can change it on the Scheduled tasks page). Marker: "scheduled …"; the agent is told to say what it scheduled in its reply.
schedule_list
No parameters. Returns every scheduled task on the Control Plane with id, name, cron, timezone, enabled, action, nextRunAt and lastRunAt.
Auto QA (end-to-end verification) runs
When Auto QA is on, Sessionboxer opens a verification run after each of your turns and asks the agent to follow the e2e-verification skill; verify opens one on the agent's own initiative. These four tools fill the run in so the Auto QA pane follows along as it happens: the cases appear when planned, each turns running, passed, failed or skipped with its note and screenshot, and the video is attached at the end. They go through the daemon's /e2e endpoint rather than the general bridge and need an open run; without one they error.
e2e_plan
Register the test cases of the current run, or skip the run when the turn changed nothing testable (an answer-only turn, research). Called once, before start_recording. Cases are numbered from 1 in the order given.
| Parameter | Type | Meaning |
|---|---|---|
cases |
array of at most 10 { title (1–200), steps (≤4,000, one per line), expected (≤2,000) }, default [] |
2 to 5 normally; up to 10 only for a very large change. Empty when skipping |
skip_reason |
string, ≤1,000 characters, optional | Why nothing is verified; no cases then |
e2e_case_start
| Parameter | Type | Meaning |
|---|---|---|
index |
integer ≥ 1, required | The case number from e2e_plan |
Marks the case running. A failed case is started again after the fix; the run keeps every attempt.
e2e_case_end
| Parameter | Type | Meaning |
|---|---|---|
index |
integer ≥ 1, required | The case |
status |
passed · failed · skipped, required |
failed means: fix the code, then e2e_case_start it again; skipped when it could not be exercised |
note |
string, ≤2,000 characters, optional | One line: what was seen, and for a failure what went wrong |
screenshot_path |
string, ≤1,000 characters, optional | A /workspace path of a screenshot of the final state; shown in the pane |
e2e_finish
Close the run after stop_recording. Cases never started are marked skipped; the run's verdict is passed when no case's last attempt failed.
| Parameter | Type | Meaning |
|---|---|---|
video_path |
string, ≤1,000 characters, optional | The recording's path as stop_recording returned it |
summary |
string, ≤4,000 characters, optional | Two or three sentences: what was verified, what failed, what was fixed |
Returns the closed run. The agent then ends its reply with a short summary that mentions the video's path, so the chat shows the player.
This chapter is generated from docs/MCP.md in the Sessionboxer repository. Found a mistake? Open an issue.