Writing

A Throwaway Windows Desktop an AI Agent Can Drive

· AI · AI Agents, Cloud

Drafted December 2025. Finished and published August 2026 as part of the migration away from WordPress.

I needed a coding agent to use a real Windows desktop: install software, operate its interface and inspect the result. A headless Windows runner could compile code but could not show the agent what an installer drew. Reusing a local VM left state behind.

I wanted the Windows desktop to behave like an ephemeral worker. One command would create and configure it, one would destroy it, and concurrent agents could each receive a clean machine of their own. The result was a small Python tool around AWS CLI.

The tool

CommandWhat it does
vm upLaunches a Windows Server 2022 instance on EC2 spot, runs the first-boot script and polls until ready. About five minutes.
vm shotScreenshots the desktop to a PNG.
vm sshOpens PowerShell over SSH.
vm downTerminates the instance. Billing stops.
driveSends one input action into the session: mouse clicks, a click on a control by its accessible name, window activation, typing, hotkeys, screenshots and resolution changes.

The instance type is set with the WINVM_INSTANCE_TYPE environment variable, which defaults to t3.large. I used that default on the spot market in London because it was enough for the job and cost about 3 to 4 pence an hour. Setting a different value launches a larger or GPU-backed Windows instance, while WINVM_ONDEMAND=1 switches from spot to on-demand capacity. The first-boot script installs no GPU driver, so a GPU instance would need that step added to the automation. The latest Windows Server 2022 AMI is fetched from SSM on every run, so there is no image to maintain.

vm up also prints a noVNC URL. I can watch the agent work in a browser and take over the mouse when necessary.

From EC2 instance to GUI worker

Launching Windows was the easy part. The first-boot script had to produce an interactive desktop that software could draw onto and the agent could control:

  • It installs OpenSSH and restricts it to key authentication.
  • It enables auto-logon so an interactive console session exists.
  • It disables UAC because elevation prompts appear on a secure desktop that injected input cannot reach. This is acceptable only because the machine is disposable and network-restricted.
  • It installs TightVNC and noVNC for supervision through a browser.
  • It writes C:\winvm\READY only after configuration finishes.

The launcher polls for that marker over SSH instead of trusting the EC2 status checks. Windows reports itself running well before the desktop, SSH server and control tools are ready. Once the marker appears, the launcher uploads the input agent and starts it as a scheduled task inside the interactive session.

SSH, RDP and VNC are restricted to the public IP that ran vm up, with the security-group rules refreshed on each launch.

How the agent drives it

The loop is turn-based: act, take a screenshot, inspect it and act again. The local drive command sends base64-encoded JSON over SSH to a small HTTP service bound to 127.0.0.1 inside the VM. That service handles screenshots, resolution changes, keyboard input and mouse input.

For standard controls it uses Windows UI Automation to find the element by accessible name and type, then clicks its actual centre. This is the relevant part of the driver:

if a == "uiaclick":
    el = find(cmd.get("name"), cmd.get("control", "any"))
    if not el:
        return {"ok": False, "error": f"no {cmd.get('control')} named {cmd.get('name')!r}"}
    r = el.BoundingRectangle
    cx, cy = (r.left + r.right) // 2, (r.top + r.bottom) // 2
    pyautogui.click(cx, cy)
    return {"ok": True, "clicked": [cx, cy], "name": el.Name}

The command drive '{"action":"uiaclick","name":"OK","control":"button"}' therefore survives window resizing and layout changes. Raw coordinate clicks remain available for software that exposes no useful accessibility information.

One desktop per agent

WINVM_SESSION selects a separate state directory containing that session’s instance record and SSH key. A second agent can launch another instance with its own desktop, filesystem and installed software. More can be added up to the AWS quota and available budget. If an installation pollutes one machine, only that disposable worker needs to be replaced.

What I would add with more time

The driver can find individual controls and return a shallow list of top-level windows, but the agent still relies on screenshots to understand most of the desktop. I would expose the full UI Automation tree as structured data, including names, roles, values, state and bounds. The agent could then inspect, invoke and wait for controls directly, using annotated screenshots and OCR only as fallbacks.

I would also give every instance an automatic expiry time, tunnel the live view and move the administrator password out of the local state file. The result would keep the same model: request an isolated Windows worker, drive it and throw it away.

← All posts