CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents | Summary
14 Jan 2026 | Paper Review Computer-Use Agents Agentic AI Safety Agent PlanningContents
- Summary
- 1 Introduction
- 2 Background
- 3 Methodology
- 4 Evaluation
- 5 Discussion
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents.
- 2026-01-14 (arXiv)
- Foerster, Hanna, Blanchard, Tom, Nikolić, Kristina, Shumailov, Ilia, Zhang, Cheng, Mullins, Robert, Papernot, Nicolas, Tramèr, Florian, Zhao, Yiren.
- University of Cambridge, University of Toronto & Vector Institute, ETH Zurich, AI Sequrity Company
- Paper
Summary
- CaMeL-NOVA adapts architectural isolation against prompt injection to GUI automation. An environment-blind Privileged Planner generates a complete branching plan before execution, while quarantined perception resolves runtime observations and coordinates within that plan.
- NOVA (Navigating via Observation, Verification, and Action) organizes plans around Observe-Verify-Act, using verify_hypothesis to select pre-written branches from runtime observations. On the matched 60-task OSWorld subset, UITars improves from 18.3% to 58.3% pass@3 relative to an unoptimized CaMeL adaptation; across three perception backends, CaMeL-NOVA reaches 65.0%–68.3% pass@5 on this selected set.
- The guarantee is control-flow integrity, not correct perception or semantic safety of every action. Cookie-mimicry and adversarial-pixel attacks demonstrate Branch Steering: manipulated observations select attacker-preferred paths or action arguments already permitted by the plan, and the tested redundancy defenses provide incomplete detection with substantial false positives.
1 Introduction
GUI actions lack the intrinsic semantics of typed APIs: click(x, y) can log in, delete data, or follow a malicious link depending on the screen. The paper separates trusted task planning from untrusted screenshots and DOM content so that environmental content cannot introduce new calls into the plan. Figure 1 shows how the interpreter obtains runtime values without returning observations to the planner.
- The P-LLM generates the plan before observing the environment; the Q-VLM supplies coordinates and state information during execution.
- The implementation is available at https://github.com/cleverhans-lab/camel-cua.
The diagram places environmental feedback inside interpreted execution rather than the planner’s context. Runtime perception supplies values for an already generated plan, preserving the set of available calls and branches while leaving their grounding dependent on potentially corrupted observations.
The feasibility argument is that enough UI contingencies can be anticipated to support useful execution, even though the full environment state space is effectively unbounded. Figure 2 contrasts short typed-API plans with larger CUA plans containing untaken recovery paths. The contributions are the measured plan-complexity gap, the CaMeL-NOVA adaptation, and the characterization of residual Branch Steering attacks.
Green nodes show one execution trace, while grey nodes represent contingencies available but not taken in that run. NOVA averages 39.7 branches compared with 11.3 in the unoptimized CUA adaptation and 3.7 in AgentDojo, illustrating the additional pre-written recovery structure used for GUI execution.
2 Background
The background contrasts reactive computer-use architectures with Dual-LLM systems and distinguishes control-flow integrity from data-flow security. Planner isolation constrains the available calls and conditional structure, while observations can still influence arguments and branch selection.
2.1 Computer Use Agents
CUAs combine perception, action, reasoning, and memory through either an end-to-end VLM or specialized agent modules. Conventional designs repeatedly adapt their plans to environmental feedback, exposing decision-making to malicious pop-ups, image patches, fine-print instructions, and other injected content.
- The defenses reviewed in the paper use attack recognition, action auditing, or browser-level policies rather than isolating the planner.
- HTTP-layer sandboxing fits browser agents, but coordinate-level interactions across desktop applications lack an equivalent general interception layer.
2.2 Dual-LLM
CaMeL emits a complete Python-like plan, whereas Fides generates successive instructions while redacting tool outputs from the planner. Control-flow integrity constrains calls and conditional structure; data-flow security additionally constrains the arguments and values those calls consume. Branch Steering attacks this second boundary by falsifying outputs of environment-querying functions such as find and verify_hypothesis.
- Strict isolation cannot solve arbitrary tasks whose required action sequence must itself be read from untrusted data.
- Earlier Dual-LLM systems use capabilities and semantic policies to restrict data flows, but transferring those policies to clicks is difficult because an action’s meaning depends on its environmental context.
3 Methodology
The attacker can modify portions of the rendered environment, including advertisements, posts, or attacker-controlled websites, but cannot access or modify the P-LLM or underlying agent infrastructure. The attacker knows the task and toolset and can approximately predict plan patterns. Attack experiments are restricted to browser-based tasks for tractability, whereas utility evaluation spans OSWorld applications.
3.1 CaMeL-NOVA: Adapting Dual-LLM to Computer Use
CaMeL-NOVA uses an environment-blind P-LLM to compile the user request into a complete plan with loops and branches covering anticipated runtime conditions. The interpreter invokes quarantined perception and binds its outputs to variables; the Q-VLM cannot rewrite the plan. The adaptation omits the original systems’ use-case-specific data-flow policies, allowing the evaluation to examine isolation without those additional protections.
Observe-Verify-Act separates state gathering, hypothesis checking, and interaction. The plan obtains screenshot or DOM observations, calls verify_hypothesis under a fixed comparison prompt, and executes a pre-written action or fallback according to the result. The main text describes a boolean verdict, while Sections B.1.2 and H express verification results using status values OK, FAIL, and UNKNOWN.
- Observation tools summarize screen content, retrieve page text or elements, and locate UI coordinates.
- Verification results become values consumed by the interpreter, not observations supplied to the P-LLM.
- Retries and alternative search strategies must already be present in the plan; execution cannot invent a missing recovery procedure.
The system prompt encodes reusable routines for browser-state checks, cookie handling, alternate finding methods, and failed-click recovery. These routines reduce task-specific planning demands without exposing runtime observations to the CaMeL planner. The reported utility therefore reflects both architectural separation and substantial engineering of the toolset and planning guidance.
3.2 Characterizing the Residual Data-Flow Attack Surface
The paper tests two probabilistic checks on quarantined perception outputs. DOM Consistency compares proposed actions and instructions against DOM structure; Multi-Modal Consensus adds an independent model that checks the screenshot, original instruction, and Q-VLM response. Either checker can halt execution, but neither provides a formal data-flow guarantee.
- DOM labels can expose visually disguised advertisement regions, but attacker-controlled HTML5 content can also change the structure communicated by the DOM.
- Independent verification is motivated by low adversarial-example transferability between diverse models; a misleading but plausible grounding result can nevertheless satisfy both checks.
Cookie Popup Prompt Injection exploits predictable consent-handling calls by placing fabricated cookie controls inside advertisements. Figure 3 shows delayed, multi-step redirection through a benign hop website. The Pixel Attack instead makes UITars return an incorrect target with a task-compatible justification; Figure 4 shows a request for “Natural Product Database” producing a click on “Rhapsido” without introducing a new call into the plan.
- The multi-step cookie variant requires repeated cookie-handling opportunities already anticipated by the plan and a hop website that does not verify its advertisement hyperlinks.
- The delayed variant places the fabricated popup on a later page the attacker predicts the agent will visit.
- The pixel experiment tests whether an incorrect grounding decision can pass verification when the returned coordinates and explanation appear compatible with the instruction.
Repeated consent-handling calls can redirect execution through multiple websites without adding a new call to the plan. A benign intermediary bypasses direct banner-hyperlink restrictions, while delaying the first fabricated popup until a predicted later page extends the attack deeper into task execution.
The adversarial banner corrupts grounding of the planner’s requested operation. The Q-VLM associates the Rhapsido link with the natural-products request and returns click(start_box='(1306,1073)') with an incorrect but task-compatible explanation, allowing the reported redundancy check to accept the action.
4 Evaluation
Utility is evaluated on OSWorld using UITars-1.5-7B, OpenCUA-32B, and Claude Sonnet 4.5 as CUA backends, with nine planner models compared on a smaller subset. Under a 15-step cap, Table 2 reports aggregate baseline success counts of 60, 76, and 109 respectively. Adding 30 tasks marked successful by default gives baseline rates of 24.39%, 28.72%, and 37.67%.
- Baseline history length is 15, except OpenCUA uses 5 because of slow throughput; maximal VLM output length is 4096 tokens.
- The published comparators cited in E.1 are UITars at 24.5 ± 1.2% pass@2 and OpenCUA at 29.7 ± 1.1% pass@3, both with 15 steps. Claude’s published 58.1% uses 50 steps, so these are not fully matched pass@1 comparisons.
- Pass@k assumes oracle selection of a successful candidate among independent attempts. It is an upper bound, not the measured performance of an implemented deployment-time selector.
- Table 2’s UITars application counts sum to 58 despite the reported total of 60; its Multi-apps count is 3, whereas Table 1 assigns 5 tasks to that category in the 60-task set.
The total row reports baseline-successful subsets of 109 tasks for Claude, 76 for OpenCUA, and 60 for UITars under the constrained execution settings. Overall rates add 30 tasks marked successful by default, making this convention essential to later improvement and retention claims. UITars’s displayed application counts sum to 58 rather than the reported total of 60.
Static plan analysis quantifies the complexity gap and how NOVA changes the generated plans. Table 3 reports 41.1 ± 1.6 tool calls and 39.7 ± 1.7 branches for CaMeL-CUA-NOVA, compared with 19.8 ± 1.7 and 11.3 ± 0.8 for the unoptimized CUA adaptation, and 4.9 ± 0.4 and 3.7 ± 0.5 for CaMeL-AgentDojo. Table 4 breaks these measurements down by application, while Figure 5 contrasts UI verification-and-recovery paths with a mostly linear typed-API workflow.
- The analysis reports N=886 plans for each CUA setup and N=242 AgentDojo plans; the reported intervals are 95% confidence intervals.
- NOVA plans average 213.3 ± 7.5 code lines, versus 71.6 ± 3.7 for the unoptimized adaptation and 51.8 ± 6.9 for AgentDojo.
- Intra-task pattern Jaccard similarity is 0.044, 0.001, and 0.393 respectively. Shared NOVA routines increase structural overlap relative to the unoptimized CUA baseline while retaining more variation than AgentDojo.
NOVA increases explicit semantic verification and action-status guards alongside plan size. Semantic verification accounts for 22.9% of its branch conditions versus 0.0% in the unoptimized CUA plans, identifying the checking structure introduced by the methodology. These descriptive statistics do not independently isolate which component causes the utility gain.
The complexity increase is visible across application suites. Chrome NOVA plans average 50.8 tool calls and 52.3 branches, while Thunderbird averages 45.9 and 48.6; the typed-API suites have smaller plans and generally greater pattern overlap. The breakdown shows that the aggregate gap is not confined to a single application.
The CUA example checks whether a browser is open and includes finding, clicking, waiting, and re-verification before proceeding. The AgentDojo example follows a short read–extract–fetch–summarize–send chain. This illustrates the additional runtime checks required to ground coordinate-level workflows within a fixed plan.
Table 1 shows a 40-percentage-point pass@3 improvement on the matched 60-task set: UITars rises from 11/60 (18.3%) without NOVA to 35/60 (58.3%) with NOVA. On this set, UITars, OpenCUA, and Claude reach 39/60 (65.0%), 40/60 (66.7%), and 41/60 (68.3%) pass@5. For the full automatically evaluable set, the table reports UITars with CaMeL-NOVA at 51/339 (15.0%) pass@1 and 77/339 (22.7%) pass@3.
- Adding the 30 default-success tasks gives 107/369 (29.0%) at pass@3, versus the undefended UITars baseline of 24.4%. This is the basis for the approximately 19% relative improvement claim, but it compares three attempts with one baseline attempt.
- Claude reaches 62/109 (56.9%) pass@5 on its baseline-successful task set. The approximately 57% retention claim is conditional on that set and is not a full-benchmark success rate.
- The similar shared-set results support a substantial planner contribution, but do not establish that perception quality is unimportant on arbitrary tasks. The NOVA ablation also changes both planning guidance and the availability of verify_hypothesis.
- Table 1’s All Tasks denominators for Chrome and VLC are 46 and 17, whereas Table 2 lists 43 and 14 automatically evaluable tasks. The aggregate denominator of 339 is retained as reported rather than reconciled by inference.
The matched UITars ablation raises pass@3 from 11/60 to 35/60 after adding NOVA’s verification primitive and planning guidance. Other columns use different baseline-successful task sets, so their percentages are conditional retention measurements rather than interchangeable full-benchmark scores. The aggregate all-task result is reported as 77/339 evaluable successes and 107/369 after default-success tasks are added, despite category-denominator inconsistencies.
Table 5 compares nine planners using UITars perception on 17 tasks: eight Chrome tasks and one task from each other application category, selected from UITars’s successful subset. GPT-5 reaches 12/17 at pass@5, followed by Grok-4 at 10/17; GPT-5.1, Gemini 3 Pro, and Gemini 2.5 Pro each reach 6/17. The authors attribute weaker results to missing fallbacks, inadequate application-navigation strategies, and limited variation across repeated plans.
- GPT-5 increases from 17.6% pass@1 to 70.6% pass@5, showing the benefit of alternative plan samples on this subset.
- The small, selected task set limits generalization of the planner ranking to OSWorld as a whole.
- Table 5 contains internal inconsistencies: GPT-OSS 120B’s application counts sum to 4 despite an overall 5/17, Kimi K2’s sum to 5 despite an overall 4/17, and Claude’s Other Apps denominator is 8 rather than 9. The review does not infer corrected values.
Planner choice changes both initial success and the benefit of repeated sampling: GPT-5 starts below Grok-4 at pass@1 but reaches 12/17 versus 10/17 at pass@5. The selected 17-task subset limits the ranking’s scope. Some application counts and denominators do not reconcile with the overall rows, so the displayed results should not be treated as a fully consistent benchmark ledger.
Fides-NOVA reaches 20/60 (33.33%) pass@1 and 40/60 (66.67%) pass@5 in Table 6, after relaxing redaction to expose selected environmental booleans to the iterative planner. The paper reports total token overheads of 1.88× for CaMeL-NOVA, 29.6× for Fides-NOVA, and 6.57× for CaMeL with Multi-Modal Consensus. Tables 8 and 9 attribute the additional expenditure to repeated planning, long plan outputs, and verifier calls.
- Fides-NOVA sets max_steps=15, max_turn=70, and max_variable_reuse=5 to cap GUI actions, total calls, and variable reuse. These bounds limit repeated queries but do not restore CaMeL’s strict environment-blind planning boundary.
- On the 17-task pass@5 cost subset, reported API expenditures are $5.40 for CaMeL, $76.07 for Fides, $11.57 for CaMeL+DOM Consistency, and $18.37 for CaMeL+Multi-Modal Consensus.
- The $0.00 undefended figure excludes locally deployed UITars computation. These are API expenditures, not complete deployment costs.
Fides uses 51,724,263 input tokens compared with 2,950,253 for CaMeL on the reported subset, showing the overhead of repeated planner calls with history. Added verification raises API expenditure further. The undefended $0.00 figure excludes local UITars computation and therefore does not imply cost-free execution.
The component breakdown assigns $3.90 to planning and $1.50 to Q-VLM functions in the CaMeL run, with verification expenditures of $6.17 for DOM Consistency and $12.97 for Multi-Modal Consensus. Planner output totals 313,061 tokens versus 133,784 for Q-VLM functions, reflecting the cost of writing extensive contingencies upfront.
Both attack classes bypass the tested redundancy checks in at least some reported experiments. Table 7 measures cookie-mimicry detection over 10 seeds: DOM Consistency flags 6/10 (60%) of standard Google-ad attacks, while Multi-Modal Consensus flags 5/10 (50%) of HTML5 attacks. Across 17 benign tasks evaluated up to pass@5, their cumulative task-level false-positive counts are 4/17 (23.53%) and 3/17 (17.65%), with both flagging 3/8 (37.50%) Chrome tasks.
- The two detection rates use different attack-defense pairings and are not a head-to-head comparison on the same attack. Static advertisement labels can expose cookie misdirection, while HTML5 banners can present cookie-like DOM elements.
- Cookie attacks can fail when plans restrict DOM searches to element types such as buttons that an attacker cannot create inside a static image advertisement.
- The pixel proof of concept uses gradient access to UI-TARS-1.5-7B and bypasses security-heavy verification. The paper does not report a broad pixel-attack success-rate benchmark across tasks or perception models.
The reported detection rates are 60% for static Google-ad attacks checked by DOM Consistency and 50% for HTML5 attacks checked by Multi-Modal Consensus; these are different attack-defense pairings. Both configurations cumulatively flag 37.50% of benign Chrome tasks by pass@5. The samples comprise 10 attack seeds and 17 benign tasks, limiting generalization beyond this narrow evaluation.
5 Discussion
CFI constrains attackers to calls already written by the planner. Redirection or denial of service may require one corrupted grounding result, whereas the paper’s sensitive-submission and exfiltration examples require compatible task-specific plans and coordinated manipulation of multiple observations. Explicit plans also provide an inspectable interface for user approval and for composing additional policy or browser-layer defenses.
- A structurally permitted click need not match the user’s intended action; translating coordinates into enforceable semantic policies remains unresolved.
- The architecture permits proprietary models to perform planning while local perception handles sensitive observations. This is a deployment option, not a measured privacy experiment.
Single-shot planning depends on anticipating initial states, navigation paths, and recovery conditions. Underspecified tasks increase branching and cost, and strict isolation cannot accommodate arbitrary action sequences whose instructions must be discovered in untrusted content. Figure 6 reports approximately 73% success by pass@20 for a Claude-based setup under oracle selection, but E.6 does not state the curve’s task-set denominator.
- OSWorld includes ambiguous task descriptions and 30 tasks without automatic evaluation; counting default successes obscures the fraction actually completed.
- Prompt engineering and state verification are implemented in the evaluated system. Task-distribution fine-tuning is proposed as a further improvement but is not evaluated.
- Independent plans can be generated in parallel because sampling does not condition on prior environmental failures. A reliable deployment-time selector or combined super-plan is not demonstrated.
The overall curve continues rising with independent attempts and reaches approximately 73% at pass@20, while application-specific curves plateau at different levels. This measures recovery under oracle selection of a successful attempt; it does not demonstrate automatic candidate selection. The accompanying section does not specify the task-set denominator.
6 Conclusion
CaMeL-NOVA demonstrates that environment-blind planning can complete useful GUI workflows when plans include extensive Observe-Verify-Act contingencies. Its security property is structural control-flow isolation under the stated threat model; the cookie and pixel demonstrations show that permitted paths and action arguments can still be manipulated.
Appendix
- A Extended Background expands the information-flow motivation. A.1 Dual-LLM and Algorithm 1 contrast static CaMeL planning with iterative Fides planning, explain capability-based data-flow policies, and identify the limitation of tasks whose action sequence is specified in untrusted data. A.2 Redundancy-based defenses connects independent checking to Algorithm-Based Fault Tolerance, N-version programming, and data diversity. A.3 How do CUAs work? details perception, action, reasoning, and memory, contrasting specialized agent frameworks with end-to-end models that maintain a reactive conversation history.
- B Extended methodology section and B.1 Combining Dual-LLM and CUA provide Algorithm 2’s single-shot generation and interpreted execution procedure. B.1.1 Threat Model explains partial environmental control and the assumptions of task, tool, and approximate plan-pattern knowledge. B.1.2 From adaptive to single shot planning describes summarize_screenshot_content, find, check_done, get_page_elements, get_page_text, find_element_by_text, and verify_hypothesis; runtime information is retained in program variables rather than used for replanning. B.1.3 Practicality of how to improve utility of single shot planning details complementary visual and DOM finding, state checks, descriptive queries, retry loops, cookie handling, and guidance to explore the current website before leaving it. B.1.4 Ablation: How to make Fides work exposes boolean results to enable conditional iterative planning, acknowledges the resulting information channel, and bounds execution with max_steps=15, max_turn=70, and max_variable_reuse=5.
- C Further Details on Redundancy Defenses describes the checks applied to quarantined outputs. C.1 Verifier Module Architecture checks outputs before they populate plan variables and halts execution when an attack is suspected. C.2 DOM Consistency compares proposed coordinates and instructions with element structure, including advertisement labels, while C.3 Multi-Modal Consensus adds an independent screenshot-based check for visual or semantic mismatches. Model diversity motivates these checks, but both remain probabilistic and add no formal guarantee beyond the core control-flow isolation.
- D Further Details on Attack Implementations and D.1 Positioning Relative to Prior Work distinguish advertisement-banner attacks from direct OS/UI injection and relate them to prior pop-up and image-patch attacks. D.2 Cookie Attack Technical Details, D.2.1 Attack Mechanism, and D.2.2 Defense Level Interaction explain fabricated consent controls in static Google images versus HTML5 advertisements: static-image attackers control an image and hyperlink, whereas HTML5 content can expose cookie-like DOM elements. D.2.3 Multi-Step Attack with Hop Website routes navigation through a benign site without banner-hyperlink verification and requires repeated cookie-handling opportunities; the long-range variant delays the first intervention until a predicted later page. D.3 Pixel Attack Technical Details and D.3.1 Objective and Rationale target an incorrect click with a plausible justification rather than conspicuous banner coordinates. D.3.2 Implementation optimizes fully masked image patches through the vision encoder, uses straight-through estimation for PNG quantization, and runs 2000 iterations on UI-TARS-1.5-7B with 2 EOT samples per iteration, Gaussian noise σ = 0.005 and clock-region noise σ = 0.05, Adam learning rate 0.8, 200-iteration cosine warm restarts until iteration 600, decay to lrmin = 8 × 10−4, and L2 gradient clipping at 10. D.3.3 Extensions and Future Work defers multi-objective, multi-model, and multi-step pixel attacks.
- E Further Utility Measurements on OSWorld collects the baseline, plan-analysis, planner, scaling, and verifier evidence. E.1 Base Utility on OSWorld reports the constrained baseline runs and explains the addition of 30 default-success tasks. E.2 Plan Analysis Methodology parses plans into AST-derived graphs with sequential and data-flow edges, classifies four types of branch conditions, and compares k=3 samples per task using sequence edit distance, cosine similarity, and pattern Jaccard similarity. E.3 Plan statistics details and Tables 3–4 quantify NOVA’s larger plans and additional verification branches. E.4 Statistics on Performance with different Planner Models supplies the selected 17-task comparison in Table 5, while E.5 Utility of UITars on Fides reports 40/60 pass@5 in Table 6. E.6 Scaling of Claude with CaMeL plots success approaching 73% by pass@20 without specifying the task-set denominator. E.7 Redundancy Defense Evaluation and Table 7 report incomplete cookie-attack detection and cumulative benign-task flags; the discussion attributes Chrome false positives to legitimate cookie banners being classified as fabricated and identifies an additional DOM-checker error on an Excel cell.
- F Token counts comparison between defenses reports pass@5 token use and API expenditure on 17 tasks in Tables 8 and 9. GPT-5 performs planning and screenshot verification, Claude Haiku 4.5 performs DOM verification, and local UITars inference is excluded from monetary totals. Long plan outputs contribute to CaMeL’s expenditure, while repeated planner calls with history account for Fides’s much larger token use.
- G Attacker Goals and Plan Structure Under CFI distinguishes redirection, arbitrary target clicks, denial of service, sensitive submissions, data exfiltration, and injection of entirely new actions. Generic UI routines can enable the first three goals. The paper’s sensitive-submission examples require a submission already present in the plan, at least two coordinated perception manipulations, and authentication at execution time; its exfiltration examples require compatible read and send steps. Calls entirely absent from the trusted plan are structurally excluded by CFI, whereas the other goals depend on the available plan structure and runtime values.
- H Example Plans includes H.1 Natural Products database example of a good plan (P-LLM: GPT-5), H.2 Natural Products database example of an insufficient plan (P-LLM: Gemini 3 Pro), and H.3 Example of a generic cookie handler snippet. The GPT-5 example contains more state checks, alternate finding methods, and recovery paths, but also permits leaving the initial website for external databases. Its final check_done result does not gate completion because both outcomes call mark_done; the Gemini example assumes a successful search state after typing and can finish without a final task-specific verification. The generic cookie snippet demonstrates the predictable, repeated grounding routine exploited by Branch Steering, so richer branching should not be equated with correct task semantics.
Brief Thoughts
The paper establishes a useful separation between fixed plan structure and fallible runtime grounding. The matched 60-task ablation measures the benefit of NOVA’s combined tool and prompt scaffolding, while selected task sets, oracle pass@k, source-table inconsistencies, and the Fides boolean relaxation limit broader comparisons. Deployment assessments should report control-flow integrity, semantic action safety, single-attempt utility, and complete execution costs separately.