Sessions
Videos, screenshots, PDFs and diagrams from the agent, in the chat
Ask the agent to record its screen, take a screenshot or export a PDF and the file plays or renders in the chat, with captions and optional narration. Markdown and Mermaid render too.
Feature page: Videos and documents in the chat — screenshots, things you can do with it and how other products compare.
Videos, images and documents from the box
Ask the agent to show you something ("record a video of the login flow", "take a screenshot of the chart", "export the report as PDF") and it saves the file in the project folder and names it in its reply. Any such file mentioned in a reply (/workspace/recordings/login.mp4, docs/report.pdf) is shown inline in the chat: videos with a player, images and SVGs as pictures, PDFs embedded, audio with controls, each with Open and Download links. The file is streamed from the box, so the session must be running to view it (a stopped one says so; Resume brings it back).
Recordings use the desktop MCP's start_recording / stop_recording tools (ffmpeg, H.264 .mp4, 15 fps by default, saved under recordings/); the agent drives the desktop as usual in between. When a recording stops it is condensed: every stretch where nothing changes on screen (a page loading, a build, the agent thinking between clicks) is cut down to a 1.5 s hold instead of being removed, so the waiting is gone but each state stays on screen long enough to read; motion plays at real speed. The tool result reports both the recorded and the final length. The agent can pass condense: false when real timing matters, or change hold_seconds.
Recordings carry captions written by the agent as it works: before each step it calls annotate_recording with a sentence about what it is doing, and each caption stays until the next one. When the recording stops the captions are burned into a band added under the desktop (so nothing on screen is covered; the band fits about three lines, and a longer caption grows upwards rather than being cut off) and also written as a WebVTT file next to the video (demo.vtt beside demo.mp4). In the chat the player lists the captions as clickable steps that follow playback (click one to jump there) and offers them as a subtitle track in the player's controls. Captions survive condensing without any adjustment: they are drawn before the static frames are dropped, so they stay attached to the frames they described, and the .vtt is re-timed the same way. The agent can pass captions: "burn", "track" or "none" to stop_recording to change what happens with them.
Captions can also be spoken: a text-to-speech model inside the box (Kokoro, no account, nothing leaves your machine) reads each caption at the moment it appears and the speech is muxed into the video as an audio track. When a sentence is longer than its step is on screen, the step's last frame is held until the sentence ends, so the video grows a little instead of the speech overrunning the next step; the .vtt and the step list are re-timed to match. Narration costs processing when the recording stops (about a quarter of the spoken time plus a re-encode: a 30 s demo with three captions takes about 4 s, one with eight about 10 s). Global settings → MCP & connectors → desktop → Narrate recordings decides: Ask when it takes longer than N seconds (the default, N = 5) narrates by itself under the threshold and otherwise delivers the silent video with the agent asking you whether to add the narration (say yes and it does, in place); Always and Never do what they say. Captions are spoken in the language they are written in when the agent passes narration_language (en, en-gb, es, fr, hi, it, pt); a voice can be chosen with narration_voice.
Markdown, diagrams and code
Replies and your prompts render as Markdown. Fenced code with a language (ts, python, bash…) is syntax-highlighted in the chat, in the rich prompt editor and in documents; a mermaid block is drawn as a diagram (flowcharts, sequence diagrams, Gantt…), with the error and the source shown if the syntax is off. Ask the agent for a document ("write the architecture to docs/arch.md with a diagram") and the .md (or .mmd) it names in its reply appears rendered in the chat, and relative links and images inside it resolve against the box's files.
This chapter is generated from docs/GUIDE.md in the Sessionboxer repository. Found a mistake? Open an issue.