Sessionboxer

Devin vs Cursor Cloud Agents vs Codex cloud vs Claude Code on the web vs OpenHands, compared (October 2026)

Six ways to run a coding agent in its own machine: where it runs, whose account pays, whether you can watch the screen, snapshot the box or switch agents.

Sessionboxer's first screen with the agent picker, repositories and the Start button

Two years ago "coding agent" meant an autocomplete that had grown ambitious. Today it means a thing that gets a task, a machine of its own and an hour, and comes back with a pull request. Six products do that well enough to compare, and they are different in ways that matter more than their model choice.

This is a comparison from the maker of one of them, so read it with that in mind. Everything below comes from each product's public documentation as of this month; the feature pages carry the same tables with more detail, and corrections are welcome as an issue.

The six, in one paragraph each

Devin (Cognition) is the original "AI software engineer": a hosted agent with its own VM, browser and, since Devin 2.2 in February, a Linux desktop it can use to test desktop apps. It bills in ACUs on Cognition's account.

Cursor Cloud Agents run Cursor's agent in a VM in Cursor's cloud. The February release gave them a full desktop and the ability to record videos of their testing; you can take the remote desktop over from the app. Part of a Cursor plan.

Codex cloud (OpenAI) runs Codex tasks in containers on OpenAI's side, started from ChatGPT, the CLI or the IDE extension, delivering PRs. Part of a ChatGPT plan.

Claude Code on the web (Anthropic) runs Claude Code in Anthropic-managed VMs, started from the browser or the Claude app, with the same Claude plan you already have.

OpenHands (formerly OpenDevin) is open source, MIT licensed, and runs each conversation in a Docker container on your own hardware, with its own agent and whatever model you have an API key for. There is also a hosted version and an enterprise one.

Sessionboxer (ours) is open source and MIT licensed. It runs on your machine or your server, one Docker container (or a Windows or macOS VM) per session, and drives the agents' own CLIs with your existing subscription: Claude Code, Codex, Cursor, Devin, OpenCode, Gemini CLI, GitHub Copilot and six more.

Where it runs and who pays

Runs on your hardware Your existing subscription Open source
Devin ✗ ✗ (ACUs) ✗
Cursor Cloud Agents ✗ ✗ (Cursor plan) ✗
Codex cloud ✗ ✗ (ChatGPT plan) ✗
Claude Code on the web ✗ ✗ (Claude plan) ✗
OpenHands ✓ API key MIT
Sessionboxer ✓ ✓ MIT

The first column is the big fork in the road. If your code may not leave your network, or your company's proxy re-signs every TLS connection, the four hosted products are out before you look at features. OpenHands and Sessionboxer run where you are; Sessionboxer also copies your machine's CA certificates into every box so HTTPS works behind Zscaler or WARP.

The second column is about money and accounts. The hosted products each come with their own billing. OpenHands calls models with an API key, which is pay-per-token. Sessionboxer uses the Claude, ChatGPT, Cursor or Devin subscription you already pay for, and shows you how much of its window you have used.

What the agent can do in its machine

Desktop with mouse and keyboard Watch live and take over VS Code and terminals inside Docker inside
Devin browser ✓ ✓ ✓
Cursor Cloud Agents ✓ ✓ — —
Codex cloud — — — —
Claude Code on the web — — — —
OpenHands browser — — —
Sessionboxer ✓ ✓ ✓ ✓

A desktop sounds like a gimmick until the agent needs to test what it built. Cursor reports that more than 30% of the PRs they merge are now written by agents working in cloud sandboxes, and the testing videos are a large part of why they trust them. Codex cloud and Claude Code on the web work in a sandbox without a documented desktop the agent drives with a mouse; they test with commands.

"Watch live and take over" is the difference between a report and a window. In Devin, Cursor and Sessionboxer you see the agent's screen as it works and can grab the mouse. In Sessionboxer the desktop is read-only while the agent works, so you do not fight over the cursor, and yours as soon as it is idle.

Things only some of them do

Snapshot and fork the machine Hand off to a different agent See the exact model calls Verify each turn on the desktop
Devin — ✗ — —
Cursor Cloud Agents — ✗ — —
Codex cloud — ✗ — —
Claude Code on the web — ✗ — —
OpenHands — — — —
Sessionboxer ✓ ✓ ✓ (Claude Code) ✓

These are the rows where running on your own hardware with the agents' own CLIs pays off. Snapshots and forks are a docker commit of the whole box after each turn; you can fork a session from any of them and try a second approach next to the first. A handoff lets Claude Code write a brief for Codex and continue the same files with a different agent. Inspect LLM shows the bytes Claude Code sent to Anthropic, which no hosted product can show without exposing its system prompt. And verification has the agent plan and run end-to-end tests on the desktop after each turn, with a captioned video. Devin and Cursor do self-testing on request after a PR; the per-turn loop is not something the others document.

Which one, then

  • You want a managed service, your code can live in a vendor's cloud and you like ACUs or credits: Devin or Cursor Cloud Agents, depending on which editor your team already uses. Both can show you the screen.
  • You already pay for ChatGPT or Claude and want to hand off small tasks from your phone: Codex cloud or Claude Code on the web. Simple and good at the small stuff.
  • Your code must stay home and you want a fully open agent with your own model key: OpenHands.
  • Your code must stay home, you already have a Claude, ChatGPT, Cursor or Devin subscription, and you want to watch the screen, snapshot the machine and switch agents mid-task: Sessionboxer. Install it with one command; the first run connects your agent in three commands.

The honest summary is that the hosted four are a product you subscribe to, and the open two are a tool you run. Both are fine answers. The wrong answer is the one that does not match where your code is allowed to be.

Features mentioned

Keep reading