Varden 1.0.1 Adds Predictive Authority to Agent Security

Varden 1.0.1 Adds Predictive Authority to Agent Security

On September 27, 2026, a Python package called Varden shipped its 1.0.0 and 1.0.1 releases on the same day. The reason for two releases in one afternoon is itself revealing: the initial 1.0.0 contained a flaw that let a protected agent disable the core security feature by writing a single JSON field. The 1.0.1 patch closed that hole within hours. But the more interesting thing is not the patch. It is what the feature does before it gets exploited: Predictive Authority, a pre-execution mechanism that performs deterministic graph reachability analysis to evaluate what dangerous capabilities become accessible if a proposed agent action succeeds.

That framing is worth sitting with. Most agent security today answers one question: Is this action allowed? Varden’s Predictive Authority asks a different one: What dangerous authority becomes reachable if this action succeeds? These are not the same question, and the gap between them is where capable multi-step agents do damage.

The Problem With Checking One Action at a Time

The standard agent security model places a gate at each tool call. An agent wants to make an HTTP request, run a shell command, read a file. A policy checks whether that agent, with its current permissions, may do that thing. If yes, the action executes. The audit log records it.

That model breaks under capable agents because a sequence of individually permitted actions can produce an outcome that no single action in the sequence would have triggered on its own. The Varden documentation describes the trajectory directly:

untrusted input
      |
credential reachable
      |
privileged service reachable
      |
irreversible action reachable

Each step might pass a conventional permission check. The agent has the right to read untrusted content. It has the right to access the credential store. It has the right to call the privileged service. A permission-per-action model sees four green lights. The trajectory ends in an irreversible external side effect caused by content the agent should never have trusted to exercise that authority.

This is not a hypothetical attack class. A September 2026 arXiv paper, Delegation Without Trust: An Empirical Gap Analysis, tested delegation security across agent frameworks at scale: 2,000 randomized scenarios over 8,100 possible actions. Bearer delegation exposed all 8,100 actions. A constrained broker approach reduced the mean reachable actions to 1.5 per authorization decision, at roughly 2.6 microseconds per decision. That gap, from 8,100 to 1.5, is what poor delegation design looks like when measured empirically. Varden’s Predictive Authority tries to close a version of that gap not at the delegation boundary but at runtime, before each action executes.

A related problem appears in Agent Memory Is a Surface for Endogenous Authorization (September 2026), which introduces EAL-Bench to test whether agent memory correctly handles permission changes, revocations, and replacements over time. Authorization errors in memory lead to unauthorized actions. The common thread across both papers: static permission checking at the tool level fails to account for how authority accumulates, flows, or propagates across a session.

How Predictive Authority Works

Predictive Authority sits in the Varden decision path after normal policy evaluation and before the audit event. The CHANGELOG is precise about its position: it is “not a second policy engine.” It runs after the existing policy has decided whether an action is permitted, then asks whether permitting that action creates a dangerous future capability path.

The mechanism performs deterministic graph reachability analysis over Varden’s accumulated authority state. The PyPI documentation states this explicitly: “It is not an LLM risk classifier. It does not attempt to guess what the agent is thinking. It does not claim to predict the next action an agent will choose.” The question it answers is fixed: “If this transition is permitted, what becomes reachable?” Not “What will the agent probably do next?”

That distinction matters for production reliability. An LLM-as-judge security layer is non-deterministic and susceptible to prompt construction. Varden’s analysis runs bounded graph traversal over a structured state model. Given the same authority graph and the same proposed action, it produces the same result every time.

Sequential Authority Accumulation

Varden builds an AuthorityState that accumulates across a sequence of actions within a session. As an agent acts, Varden updates what authority the agent currently holds and what it can reach from that position.

Five built-in Predictive Authority policy classifiers identify distinct hazardous trajectory patterns:

  • credential_acquired: An action brings a credential into the agent’s reachable scope.
  • untrusted_to_external_path: Untrusted content is on a path toward an external destination.
  • cross_trust_domain_path: A reachable path crosses a trust domain boundary.
  • irreversible_action_reachable: An irreversible operation becomes reachable from the current authority state.
  • authority_expands: The proposed action increases the agent’s authority graph.

These classifiers fire as trajectory patterns emerge over multiple actions, not only at terminal actions. A credential acquisition triggers credential_acquired before the agent uses that credential downstream. The dashboard surfaces the trajectory forming, not just its outcome.

Bounded Analysis and Truncated Results

The reachability analysis uses independent limits: graph size, traversal depth, per-node visits, and path count. This prevents an adversarial or unusually large authority graph from creating unbounded computation in the control plane.

Truncation behavior is specific. If analysis hits a limit and reports TRUNCATED, any hazardous paths found before truncation remain valid. The design rule, stated directly in the documentation: “Incomplete analysis is never converted into a safe result.” A credential reachability path found on the first 50 nodes holds even if the remaining 1,000 nodes were never examined. A TRUNCATED result containing a hazardous path is not an ambiguous result. It is a confirmed finding with incomplete coverage of the remainder.

Varden also includes protection against a specific evasion pattern the documentation calls “horizon-camping,” where an attacker structures an authority graph so that dangerous paths appear just beyond the configured traversal depth. The bounded analysis addresses this at the traversal boundary.

Observe Mode and Enforce Mode

Predictive Authority is opt-in and provides two operating modes. In observe mode, Varden records its recommendation and supporting evidence without altering the existing policy decision. An operator can run observe mode for a week, accumulate trajectory data, and understand what Predictive Authority would have done before committing to enforcement.

In enforce mode, Varden may strengthen an existing policy decision. It can only move a decision in one direction:

allow  -->  require_approval
allow  -->  block
warn   -->  require_approval

block  -->  allow  (never)

This makes Predictive Authority additive rather than competing with existing policy. A policy that already blocks a tool call is not overridden by PA analysis. Configuration for Predictive Authority comes only from the policy file’s predictive_authority section and VARDEN_PA_* environment variables. Agents cannot influence it, which is relevant to the patch.

What the 1.0.1 Security Patch Fixed

Three security issues were closed in 1.0.1, and all three share the same root cause: the control plane trusted client-supplied input it should have rejected.

The most significant: in 1.0.0, Varden read Predictive Authority configuration from action.metadata.predictive_authority_config, a field the protected agent writes. Sending {"mode": "off"} in that field turned off operator-enabled enforcement for that agent’s actions. An agent could disable the security layer by including a specific metadata key in its own action payload. The fix: configuration now comes only from the policy file and environment. Client-supplied predictive_authority* metadata is stripped at ingest.

The second issue: /sdk/guard and /sdk/log used the tenant from the client payload rather than the credential’s tenant, creating a tenant isolation gap. The third: POST /predictive/demo reset the process-wide session registry. Any viewer with access to that endpoint could discard every agent’s accumulated authority state. The demo endpoint now runs in an isolated store.

The release notes recommend upgrading if running Predictive Authority in enforce mode. The 1.0.0 flaw was exploitable by the protected agent itself, not only by an external attacker. In systems where prompt injection is a realistic vector, that is a meaningful threat model distinction.

Validation and Scope

The 1.0 release tree passed 726 tests total: 692 Python tests, 26 frontend tests, and 8 browser smoke tests. Predictive Authority includes deterministic, bounded, and adversarial validation as separate test categories. These are project-reported numbers, not externally audited.

Varden covers HTTP (requests, httpx, urllib), subprocess execution, filesystem operations, MCP via a gateway, and browser interactions via its Web Shield component. The runtime coverage model explicitly categorizes each surface as ENFORCED, PARTIAL, NOT_ROUTED, or UNCOVERED. Filesystem coverage remains PARTIAL due to documented residual TOCTOU gaps. Raw sockets, aiohttp direct, and urllib3 direct paths are UNCOVERED. Varden reports these gaps rather than asserting coverage it does not have.

Limitations and Open Questions

Varden labels Predictive Authority experimental in 1.0.x. Several limitations matter for production evaluation.

The multi-worker constraint is architecturally significant. Predictive Authority holds live session state per control-plane process. Running multiple uvicorn workers splits a session’s action history across processes. A trajectory that crosses process boundaries is invisible to both workers. The documentation is direct: run a single control-plane worker when Predictive Authority is enabled. For high-throughput deployments, this is a real constraint that requires either vertical scaling of the control plane or accepting that cross-worker trajectories are missed.

Session eviction also matters. The session registry is bounded by VARDEN_PA_MAX_SESSIONS (default 10,000) and VARDEN_PA_SESSION_IDLE_SECONDS (default 24 hours). An evicted session restarts with a fresh authority state. A dangerous trajectory that spans an eviction boundary loses its prior context. The documentation notes that evictions are reported in registry stats, so sizing the cap for the actual agent workload is necessary configuration.

Varden is not an OS sandbox and does not completely mediate arbitrary Python execution. An agent using a saved pre-patch function reference, an unsupported network client, or a non-routed MCP server can bypass instrumentation. The provenance model for cross-server MCP operates on the “supported host path” via trace/session provenance keyed by trace_id. MCP calls not routed through the Varden gateway are not tracked, meaning an agent can acquire authority through a non-routed MCP path without Varden seeing the acquisition.

What This Means for Engineering Teams

For teams running AI agents with meaningful tool privileges, the architectural implication is that permission checking at the tool level is necessary but not sufficient. An agent that legitimately holds credentials, shell access, and outbound HTTP rights can produce a dangerous outcome through a sequence of permitted operations. Tracking what authority state accumulates over a session, and what becomes reachable as it grows, is a separate problem from checking individual action permissions.

The provenance model Varden introduces is worth understanding independently of the product. Varden distinguishes between authority the agent holds and authority that the information currently influencing the agent is authorized to exercise. An agent reading a malicious GitHub issue still holds deploy credentials. The issue content does not hold those credentials. A security design that asks “is the current influence source authorized to trigger this action class?” is asking a different question than “does the agent have permission to perform this action?” Both questions matter. Only one gets asked by most current systems.

For teams building agentic automation workflows, the five PA classifiers serve as a useful framework for auditing existing action sequences. For each workflow, the practical questions are: does any action in this sequence acquire a credential? Does any step cross a trust boundary? Could the sequence reach an irreversible action through individually permitted steps? Answering those questions identifies where trajectory-level controls add value, regardless of whether Varden is in the stack. For more on how agent runtime architecture affects security surface, the post on optimizing the agent inference stack covers related tradeoffs.

Practically, the recommended starting point is observe mode for several days. Collect the trajectory evidence, identify which patterns in the existing workflow trigger PA recommendations, understand the false positive rate before committing to enforcement. Only after that analysis does enable mode make sense, and only with a single control-plane worker until the multi-worker limitation is resolved in a future release.

Key Takeaways

  • Varden 1.0.1 ships Predictive Authority: deterministic graph reachability over accumulated authority state, running before each agent action executes, not only at terminal actions.
  • In 1.0.0, protected agents could disable Predictive Authority enforcement by sending {"mode": "off"} in action metadata. Version 1.0.1 strips client-supplied predictive_authority* metadata at ingest.
  • In enforce mode, Predictive Authority can only strengthen existing policy decisions, never weaken them. An existing block cannot be overridden.
  • Five built-in classifiers surface trajectory patterns: credential acquisition, untrusted-to-external paths, cross-trust-domain paths, reachable irreversible actions, and authority expansion.
  • Truncated analysis is never treated as a safe result. Hazardous paths found before the traversal limit remain valid regardless of whether the full graph was analyzed.
  • Predictive Authority requires a single control-plane worker. Multi-worker deployments split trajectory history across processes and miss cross-worker chains.
  • The package is Apache-2.0 and installable via pip install varden==1.0.1. A demo, dashboard, and policy packs are included for evaluation without production configuration.

Work With Origins AI

Origins AI builds production AI systems for engineering teams. If your agents hold meaningful tool privileges and you need to reason about authority trajectories before they become incidents, talk to our team.