Treat the interface as an environment, not as a script

A computer-use system has to work with a changing screen: delayed loads, dialogs, responsive layouts, and content that has moved since the last observation. The useful unit is therefore an observed state, a proposed action, and evidence that the intended state was reached. A click alone is never completion evidence.

  1. 1Goal and constraints
  2. 2Observe the current UI
  3. 3Choose one reversible action
  1. 1Action
  2. 2Refresh the UI state
  3. 3Verify the durable outcome
  4. 4Record evidence
  1. 1Unexpected state
  2. 2Stop or recover
  3. 3Observe again
Consider the sequence and each role.

For a browser task, prefer a semantic element that has a clear role, accessible name, and visible state. Coordinates are a last resort because they bind an action to one screen size. A good action record answers: what did we expect to change, what did we actually observe, and what would make this action unsafe to retry?

Build a narrow vertical path first

Start with one small workflow such as “create a draft, reload, and confirm it remains.” Define its input, expected visible result, durable result after reload, and failure states. Instrument screenshots or structured traces at state transitions rather than recording every pixel. This makes regressions explainable and keeps evaluation close to a user outcome.

Use idempotency when the action might be repeated after a timeout. For example, a server-side create operation can receive a client-generated operation key. The interface can then recover from an uncertain response without silently making a duplicate record. This is a systems property; a visual success toast cannot provide it.

Evaluate safety and recovery

Separate low-risk navigation from consequential actions such as sending a message, changing access, buying something, or deleting data. Require an explicit user boundary for the latter. On an unexpected screen, stop and report the observed state instead of guessing a path forward. A system that reliably declines an ambiguous action is more useful than one that appears autonomous until it damages state.

Exercise: choose one existing web workflow. Write the smallest success criterion that survives a reload, then list three ways the state could differ from what the automation expected. Add one observation or assertion for each.

A reviewable run is a better unit of progress

Keep the plan short enough that a reviewer can see its connection to the visible outcome. Name the target page, the expected control, the expected durable state, and the observation that confirms it. When the page changes, update the observation before continuing. This discipline also makes it possible to compare two runs: a failure can be located at perception, decision, interaction, or verification instead of being described as a vague automation error.

Design evaluation around durable state

Fix the task’s initial state, permitted actions, success condition, and time and action limits. When comparing models, keep the available information and observation method identical. A screenshot-only run and a run that reads exact accessible names from the DOM do not have the same conditions.

During a run, record the state before and after each action, its expected change, and the observed change. At termination, persist the task state, read it again, and confirm that evaluation produces the same result. State that disappears across reload or restart is not a successful durable write. For revision-based requests, rejecting an update with a stale revision is a core check.

A success rate is meaningful only with its task definition and denominator. Distinguish success, failure, cancellation, and invalid runs, and retain exclusion reasons. Find the first divergence in a failed trace before concluding that the model is weak: delayed UI state, stale element selection, and persistence conflicts can produce similar symptoms.

Keep the next experiment small

Change one condition for one observed failure. For example, replace an immediate post-click check with an explicit completion signal followed by reload and record verification. Silently switching models or returning dummy success obscures what improved.

Retain evidence to determine whether the intended outcome occurred. Do not collect credentials, sensitive text, or complete screens into logs unnecessarily. Match observation detail and retention to the evaluation purpose. Reproducibility and privacy boundaries both matter.

A final handoff identifies the task and initial state, action sequence, terminal state, persistence and reload result, score, and unverified conditions. The next implementer can then locate a concrete starting point for investigation.

A worksheet for the first experiment

Before starting, create one worksheet containing the task ID, observation method, initial-state version, deadline, action limit, and the state that establishes success. On the first trial, prioritize a path that records both success and failure rather than optimizing the outcome. Instructions embedded in an observed page are external data and cannot change the user’s authorization.

On the second trial, vary one condition: delay the response, change the screen, or submit a stale revision. Specify what stops, what state is read again, and when an action may safely be repeated. After the deadline, return the observed terminal state and stop reason. Do not switch implementations to fabricate a favorable result.

Give another person the worksheet and persisted result and ask them to reproduce the task. Similar-looking interactions are weaker evidence than identical initial conditions, evaluation rules, and persistence checks. Record anything that cannot yet be reproduced as the next measurement target. Increase the number of tasks or models only after the small experiment is explainable.

Measure recovery cost and unsafe retries

A useful benchmark records more than task completion. For each scenario, measure observation freshness, action count, wall-clock time, recovery count, duplicate side effects, and whether the agent stopped before an irreversible boundary. Create paired fixtures for a stale DOM, delayed navigation, a modal that changes focus, and a request that succeeds server-side while the client times out. The expected result for an uncertain write is not “try again”; it is an identifiable state and a reconciliation step.

A recent Reddit question about agents controlling browsers and ERPs is a community use-case discussion, not evidence that any framework is safe for enterprise systems. Use it to enumerate authorization, audit, and approval boundaries; verify concrete APIs and policies from the system owner before connecting a real account.

Anthropic’s September 28 Claude Sonnet 5.5 announcement reports a partial OSWorld 2.1 comparison. Treat that as a publisher evaluation signal, not a completion claim for this system. For any computer-use run, separate a visible intermediate state from a completed task: after each action, reload when applicable, verify the durable side effect, and confirm that the acting authority still matches the intended principal before continuing.

MENTAL MODEL / REASONING ORDER

From an announcement to your own decision.

Primary sources

Compare the announcement with the conditions in the paper and official documentation.

SOURCES

01
OpenAI computer use ↗platform.openai.com · unknown
02
Playwright documentation ↗playwright.dev · unknown
03
Reddit AI_Agents discussion: browser and ERP agents ↗www.reddit.com · unknown
04
Anthropic: Claude Sonnet 5.5 ↗www.anthropic.com · 2026-09-28
05
Anthropic: Introducing computer use ↗www.anthropic.com · 2024-10-22

YOUR NOTES