Gorio Tech Blog search

Summary explanation: Agent Safety Should Be a Runtime Contract

|

Contents

This article explains the key points of Agent Safety Should Be a Runtime Contract.

  • 2026-08-11 (arXiv), Preprint
  • Ng, Albus W., Han, Yi, Zhang, Jusheng, Wang, Wenhao.
  • Vast Intelligence Lab, Southwest University, Sun Yat-sen University
  • Paper

Read this article in Korean


Summary

  • This position paper argues that safety for consequential AI agents should be treated as a runtime contract enforced by the harness, not solely as a property acquired during training. The proposed contract combines preventive controls that constrain dangerous actions with evidential gates that accept completion only when a trajectory contains verifiable artifacts; the argument draws on incident, false-completion, trajectory-schema, and proceedings-title audits. For harness, see this post. For harness, see this post.

1 Introduction

The paper defines a harness as non-model inference infrastructure connecting a foundation model to the world, including input sanitization, output filters, permissions, sandboxes, oversight, and execution tracing. A trajectory records observable operations and effects, while evidence-gated submission refuses to accept a task as complete unless specified artifacts can be verified.

2 Why Model-Only Alignment Fails on Both Counts

The authors contrast model-only alignment with two historical deployment practices. Computer security developed defense in depth rather than relying on a correct trusted component, and experimental science shifted acceptance from personal assertion toward externally checkable records; the paper presents runtime prevention and evidence as analogous requirements for agents.

3 Mismatches Between Model Alignment and Agentic Deployment

The paper identifies five mismatches between model alignment and consequential deployment: statistical proxies versus formal specifications, finite training distributions versus open-world use, unverifiable internal reasoning versus replayable trajectories, plausible output versus grounded evidence, and a single defensive layer versus compositional defenses. It argues that these are deployment-level gaps rather than problems resolved merely by scaling the underlying model.

3.1 Preventive Mismatch

On prevention, learned reward models are characterized as statistical proxies that can be optimized in ways that diverge from intended preferences. Permission rules, tool whitelists, sandboxes, and resource limits instead provide explicit runtime constraints, though the paper does not claim that formal rules are immune to manipulation.

3.2 Evidential Mismatch

For evidence, the paper treats model self-reports and model-generated reasoning as soft evidence because their reliability depends on the model. It recommends externally checkable artifacts such as replayable logs, test reruns, citation lookups, diffs, and snapshots, which can be assessed without inspecting or trusting the agent’s internal state.

3.3 Combined Mismatch

The combined mismatch is that model-level alignment can become a single point of failure under jailbreaks, fine-tuning degradation, or distribution shift. The proposed response layers input filtering, tool gating, output screening, and sandboxing with trajectory monitoring, evidence gates, and approval gates so that defeating one mechanism does not by itself establish safety or completion.

4 The Two Faces of the Safety Harness

The safety harness has a preventive face that governs execution and an evidential face that determines whether a claimed task outcome is accepted. The paper connects them through a compositional-gating proposition, framing both as parts of one runtime contract.

4.1 The Preventive Face: Mechanism Taxonomy and Design Principles

Preventive mechanisms are organized by timing into preventive, detective, corrective, and structural categories. The design principles adapt defense in depth, least privilege, fail-safe defaults, complete mediation, and auditability: risky capabilities are denied by default, model-to-world interactions are mediated, and actions are recorded in tamper-evident form.

4.2 The Evidential Face: Agent Trajectory Schema and Evidence Chain

An Agent Trajectory is defined as a finite hash-chained sequence of typed events with timestamps and payloads. Hard evidence is provided by an event that a deterministic verifier can assess against external reference state, whereas soft evidence relies on the correctness of model-generated content; an Evidence Chain must provide hard evidence for every requirement in a task-specific schema.

4.3 Compositional Gating

Each preventive layer is modeled as a deterministic finite-state monitor, and each evidential gate as an evidence-chain checker. With pairwise disjoint monitor alphabets and independent verifier sets, parallel composition enforces the conjunction of safety properties and accepts completion only when all required evidence chains verify; shared events require assume-guarantee reasoning, and general verification may be exponential.

The diagram places the agent's reasoning, planning, and tool use between environmental access and two harness paths. The preventive path filters inputs and mediates permissions, tools, and sandboxed execution, while the evidential path gathers logs and hashes, verifies artifacts, and gates whether an output is marked verified.

Architecture of a two-faced AI-agent harness that combines execution controls with evidence-gated acceptance.
Architecture of a two-faced AI-agent harness that combines execution controls with evidence-gated acceptance.

5 Empirical Evidence

The four audits are descriptive and counterfactual rather than causal evaluations. Of 52 incident-survey cases, 40 are coded fully preventable by a functional harness layer, 11 partially mitigable, and one primarily an internal-goal alignment case; the false-completion audit contains 31 non-contested core cases plus one disputed illustrative case. The trajectory audit finds submission-like evidence gates documented in 2 of 12 systems, and the title-level audit of 28,560 NeurIPS, ICML, and ICLR papers from 2023–2025 estimates a pooled 8–12× training-time versus deployment-time imbalance, subject to documented public-record and truncation limitations.

The table aggregates the paper's four descriptive audits: 52 incident cases, 31 non-contested false-completion cases plus one disputed illustration, 12 public systems or harnesses, and 28,560 accepted conference papers. It reports the headline counts supporting the authors' claim that documented deployment failures coexist with sparse submission gates and publication attention concentrated on training-time interventions.

Headline results from the incident, false-completion, trajectory, and proceedings audits.
Headline results from the incident, false-completion, trajectory, and proceedings audits.

The matrix compares twelve systems or harnesses across structured logs, test runs, file diffs, tool outputs, screenshots, and submission gates. It reports that 11 systems document tool outputs, 9 document file diffs, and only 2 document submission-like gates; the text distinguishes GitHub Copilot's PR/CI workflow from OSWorld, which is a benchmark harness rather than a deployed product.

Evidence-gating capabilities documented for twelve public agent systems and harnesses.
Evidence-gating capabilities documented for twelve public agent systems and harnesses.

6 Example: Code-Patch Submission

For a coding agent submitting a bug-fix patch, the paper instantiates the two-faced contract. The preventive side limits execution and permissions, while the evidential side requires a linked patch diff, a timestamped developer-test invocation and result, and a commit hash before the harness accepts completion.

6.1 Preventive Face

The example uses five preventive layers: a Docker sandbox without network access and restricted to the project root; a tool whitelist that requires approval for rm, git push, and curl; filesystem scope guards; a monitor for credential-read-then-write patterns; and rollback with human escalation. These mechanisms are intended to restrict or contain side effects before they persist.

6.2 Evidential Face

The code-patch evidence chain contains file_write, shell_exec, tool_result, and commit events linked by the trajectory hash chain. Its submission schema requires a commit, a zero test exit code for that commit, and a non-empty diff; the authors argue that fabricated results without a matching trace fail the chain, even if preventive controls have been overridden.

7 Counterarguments

The paper argues that runtime controls complement rather than replace model alignment. It limits evidence-gating to effects with checkable acceptance conditions, recommends human approval for non-idempotent actions in open-ended settings, and acknowledges the engineering cost of defining task-level schemas.

8 Conclusion

The conclusion calls for a shared runtime discipline built around canonical trajectory schemas, task-specific evidence requirements, and public failure reporting. Its central architectural claim is that the relevant safety unit is a trajectory with checkable evidence rather than the model alone.

Appendix

  • The appendix limits the contract to actions and submissions rather than goals, notes that polynomial compositional verification depends on verifier independence, and recognizes that classifier-based components retain narrower-domain fragility. It also identifies English-language oversampling and the current restriction of evidence schemas to tasks with established correctness criteria, then proposes canonical hash-chained schemas, published task schemas and differential verifiers, system-level benchmarks, and audit-grade logging.

Brief Thoughts

The paper usefully separates preventing unsafe side effects from verifying that an agent completed its claimed work, and the code-patch example makes that distinction concrete. Its empirical material is primarily descriptive and counterfactual, particularly the incident coding and proceedings-title audit; controlled studies measuring the effects of specific gates on safety, task success, latency, and operator workload would substantiate the systems proposal further.