Verification agents and screen recordings
Every Trevize run stops at a pull request, and until now the evidence attached to it was a test report. Tests tell you the code does what the tests say. They don’t tell you the button works. So runs now open the product, use the change, and record it. When one gets stuck on a login it can’t get past, you can watch its screen and take over.
Proof that the change works
Feature, Quick Change, and Bug Fix now take an optional Verify step that runs after tests. Turn it on in the blueprint and every run on it ends with a demonstration instead of a claim.
The agent starts the product on the surface the change actually touched, whether that’s the web app, a CLI, a desktop build, or a mobile app, and uses the change the way you would. Then it records itself doing it. Instead of an agent telling you the signup flow works, you get a short video of the signup flow working: the form filling in, the error not appearing, the page loading.
The clip plays inline in the reply and you can seek through it, so you can jump to the ten seconds that matter instead of watching the whole thing. Playback stays private to your workspace.
Everything the step produced lands on the pull request under a single Verification section, screenshots and recordings included. One place to look, instead of scattered comments.
Watch the agent’s screen, and take the wheel
Sessions and autonomous tasks now have a screen. Open it and you get a live, view-only desktop of whatever the agent is looking at: the browser it opened, the app it’s clicking through, the dialog it’s stuck on. Runs used to go quiet behind a wall of tool calls. Now you can just look.
When watching isn’t enough, press Take Control and the desktop is yours. You drive with your own mouse and keyboard, in the same session the agent is using. This is the fix for runs that used to simply stall: a login wall the agent has no credentials for, a two-factor prompt, a consent screen it can’t get past, a payment form nobody wants an agent filling in. Sign in yourself, clear the obstacle, hand control back, and the run continues from there. Before this, your options were to abandon the run or write a follow-up describing what to click and hope for the best.
Added
- Docker Engine and Compose are available in the agent’s environment. The daemon stays off until Docker-backed work needs it.
- Trevize-managed agent instructions now live in
.trevize/AGENTS.md, so Trevize stops editing your repository’s own instruction files. trevize secrets list,set, anddeletemanage workspace secrets from the CLI.
Improved
- A finished implementation only moves to a pull request on passing tests. “Looks done” is no longer treated as evidence.
- Work that completes but is blocked on a remote action it lacks permission for stays open as Needs attention instead of quietly closing.
- Change models mid-run. Pick a different model with your next follow-up and the run continues, resuming the session you had inside the same harness or starting a fresh one on the same task when you switch harness.
- Rate-limited or temporarily unavailable steps continue on a different model right away, without pausing the whole workflow first.
- Completed tasks stay in the active list after their pull request is merged, until they archive normally.
- The test step no longer pads the suite with snapshot-only or coverage-gap tests, and removes agent-added tests that fall below that bar.
- Failed Blueprint workflows restart on the controller and workspace they were already using when that’s still possible, and only get a clean replacement when the workspace can’t be reused.
- Resume on a paused task is acknowledged immediately, so you can tell the click registered.
- GitHub App installations are resolved per repository and refreshed right after you reconnect GitHub.
Fixed
- Verify inspects the committed branch diff against the base branch, not a bare working-tree diff.
- Manually opened pull requests stay linked to their task, including after a partial failure while opening the PR.
- Fixture and log pull request URLs no longer attach a fake PR to a task.
- Step outcomes the workflow can’t route are answered with a correction instead of crashing the run.
- A successful retry after an attention hold continues the workflow instead of pausing again.
- Paused, provider-limited tasks resume from a model-selection retry without a second action.
- Recovered Claude Code and OpenCode steps no longer start a duplicate agent.
- OpenCode sessions interrupted by a deploy get an explicit continuation instead of running without an agent.
- Runs no longer stay marked active and healthy after their controller has entered a definitive error state.
- The production worker no longer crashes on boot while loading sharp.
As always, everything above shipped through Trevize’s own runs, verification recordings included.