AlloyDB Gives AI Agents Their Own Isolated Database Compute Without Touching Production
Google Cloud published a detailed architectural description of AlloyDB’s new agentic database architecture on September 25, 2026, authored by Amit Ganesh and Sailesh Krishnamurthy, both VPs of Engineering for Databases. The post announces a preview of “PostgreSQL for agents” in AlloyDB, a system where AI agent fleets spin up thousands of independent, full-featured PostgreSQL database nodes in seconds, query live production data with sub-second freshness, and leave the production cluster completely undisturbed. The architecture is available in preview now, with early access via a sign-up form.
The Problem With Every Existing Database Architecture
Ganesh and Krishnamurthy open with a pointed observation: every major database scaling architecture of the last four decades makes at least one structural compromise on the three properties that agentic workloads require simultaneously. Traditional independent replicas (shared-nothing storage) achieve isolation and sub-millisecond latency, but scaling requires rehydrating gigabytes or terabytes of storage per replica, measured in hours, not seconds, and completely incompatible with agent reasoning bursts that complete in under a minute. Disaggregated shared-storage servers (like Aurora’s log-offloading approach) scale compute quickly and achieve low latency, but agent I/O and production I/O draw from the same storage pool, creating shared fate. Object storage with shared block servers (the emerging Snowflake-style architecture) can serve hot data at low latency from the cache tier, but a block server cache miss falls back to object storage at tens of milliseconds, an order of magnitude higher than traditional database storage, and the shared block servers create the same isolation problem as shared-storage servers.
The authors tested one commercially available service in the third category by running concurrent index lookups over a dataset larger than available DRAM, then scaling from one to eight read replicas. Adding replicas provided less than 2× aggregate throughput increase before saturating. Primary database throughput dropped by more than 75% as replicas were added. This is not a configuration problem, it is a structural consequence of shared physical infrastructure.
Agentic workloads break all three constraints simultaneously. They cannot be vetted in advance (so pre-provisioning fails). They generate large, bursty, unpredictable query volumes (so shared infrastructure fails). They require low-latency access with the full power of the database engine, indexes, vector search, full-text, spatial, SQL, not just row scans (so direct object storage access for agents fails). And they must not affect the production systems running the actual business.
The Three Tenets and How AlloyDB Satisfies All of Them
The AlloyDB agentic architecture is organized around three principles that must hold simultaneously:
Isolation: agents must read live production data with sub-second freshness over a data path that shares no database components with the primary cluster. This is physical separation, not a quota or a rate limit.
Latency: sub-millisecond baseline storage I/O, with no performance cliff on cold cache misses.
Scale: zero to thousands of compute nodes in seconds, auto-scaling down to zero when agents finish, billed per second of activity.
AlloyDB achieves all three through vertical integration across Google’s storage, network, and compute layers.
Storage: Colossus as the Foundation
Colossus is Google’s exabyte-scale distributed storage system that underpins Search, YouTube, Gmail, Drive, Spanner, and BigTable. Ganesh and Krishnamurthy highlight three properties of Colossus that make the AlloyDB architecture possible.
First, direct sub-millisecond I/O. A database node opening a Colossus storage stream receives a handle describing where data physically resides. Authorization and metadata resolution happen once, when the stream is created; every subsequent read goes directly to the disks over an optimized network protocol. There is no cache tier to warm, no intermediary to miss. Every storage read completes at sub-millisecond latency across the full dataset.
Second, massive aggregate throughput. A single AlloyDB database can receive up to 15 TB/s of aggregate throughput and 20 million queries per second from Colossus, with no need to provision bandwidth and with an unlimited number of concurrent hosts. A fleet of 1,000 agent nodes is, in the authors’ words, “a fraction of the load the storage already serves.”
Third, physical segment partitioning. AlloyDB serves agents from a dedicated set of Colossus segments that are separate from the production data path. Agent I/O is deliberately spread away from production I/O rather than contending with it. No component, compute, network, or storage, is shared between agent workloads and the production cluster.
Network: Jupiter
Compute and storage are connected through Jupiter, Google’s high-capacity data center network. A single Jupiter fabric connects more than 100,000 servers with 13 petabits per second of bisection bandwidth. Because Jupiter provides high bisection bandwidth with predictable low latency across the fabric, agent nodes can be scheduled anywhere in the cluster with consistent storage access. As the agent pool scales from zero to thousands, the underlying interconnect absorbs the expanding traffic without creating placement bottlenecks.
Compute: Ephemeral microVM PostgreSQL Nodes
At the compute layer, agents connect to the AlloyDB agent pool through the Model Context Protocol (MCP). The pool consists of AlloyDB agent nodes, each a fully functional PostgreSQL database engine running inside a lightweight, isolated microVM, with read-only access to the up-to-the-second state of the database. These instances are isolated from each other and from the primary cluster. They provision in response to agent requests and stop automatically when agents finish. Billing is per second of agent node activity; a burst that uses 1,000 nodes for tens of seconds is charged only for the resources consumed.
Agent nodes access every capability of the full PostgreSQL engine: point lookups, index traversals, vector search, full-text search, spatial search, columnar scans, and federated queries across the lakehouse via BigQuery and Spark. Unlike RAG architectures that create stale copies of data, agents query the live production state directly, with freshness measured in under one second.
Benchmark Results
Google tested AlloyDB by running concurrent index lookups over a dataset larger than available DRAM, scaling from one agent node to 1,000.
- Throughput scaled near-linearly from 3,900 QPS at one node to 41,000 QPS at 10 nodes, then maintained near-linear scaling to 1,000 nodes.
- Aggregate throughput reached 3 million QPS across 1,000 compute nodes, driving over 8 million IOPS in Colossus.
- Scaling from 1 to 1,000 agent nodes produced no measurable impact on primary cluster performance, zero primary degradation.
In a second benchmark running concurrent full table scans across 2,100 agent nodes, aggregate scan throughput exceeded 1 terabit per second. For comparison, the competitor test with 8 read replicas achieved less than 2× throughput increase while dropping primary throughput by more than 75%.
Limitations and Open Questions
These benchmarks run on Google’s own infrastructure stack. The combination of Colossus’s sub-millisecond storage latency, 15 TB/s aggregate throughput, and physical segment partitioning, plus Jupiter’s 13 petabit/s network, is not available outside Google Cloud. Teams evaluating AlloyDB’s agentic architecture are evaluating a vertically integrated system where the database, storage, and network were designed together over decades. The economics and performance characteristics of architecturally similar designs built on different infrastructure will differ.
Agent nodes are currently read-only. This is sufficient for the primary use case, agents reasoning over live operational data, but limits agent write patterns to going through the production cluster via normal application paths. The announcement does not describe a timeline for read-write agent access.
The architecture is in preview. The announcement provides a form to request early access and links to documentation. The pricing model (per second of agent node activity) is described conceptually; specific pricing is not published in this announcement.
For teams building large agent populations, the latency of individual agent-node spin-up (described as “seconds”) is relevant. A reasoning burst that requires a new node on each step may pay spin-up latency on each iteration. The announcement does not characterize this cost in detail.
What This Means for Engineering Teams
The most consequential implication of this architecture is that database compute can become part of the agent runtime rather than a separately provisioned service. Today, most architectures give agents access to data through one of three mechanisms: shared connection pools to the production database (risking production impact), pre-built vector or RAG copies (stale and semantically limited), or dedicated read replicas (slow to provision, expensive to maintain at scale). AlloyDB’s agent pool offers a fourth option: full-featured PostgreSQL compute that spins up with the agent, queries the same live data as production, and disappears when the agent finishes, leaving nothing running and incurring no ongoing cost.
This matters most for teams building agents that need to perform complex database reasoning, joining tables, running aggregations, executing vector search alongside relational queries, rather than simple point lookups. RAG with a vector store answers “what documents are semantically similar to this query?” A full PostgreSQL agent node can answer “which customers in region X who ordered products Y and Z in the last 30 days have average order value above $500 and have not been contacted this quarter?”, and do so against live production data without any extract or copy step.
For teams building production AI agents that require data access, this architecture changes the cost structure of database isolation. Instead of paying for dedicated read replicas 24/7 to be ready for agent queries, teams pay for compute only during agent reasoning. An agent fleet that runs for 10 seconds per query across 1,000 concurrent sessions uses 1,000 node-seconds of compute, then releases everything. The engineering implication is that the database tier can now scale in exactly the same way as the compute tier: on demand, per task, and at agent speed.
Key Takeaways
- AlloyDB’s agentic architecture provisions ephemeral microVM PostgreSQL nodes that read from dedicated Colossus storage segments, physically separate from the production data path, no shared components at compute, network, or storage layers.
- In testing: scaling from 1 to 1,000 agent nodes produced zero measurable primary cluster degradation and 773× throughput scaling to 3 million QPS and 8 million IOPS in Colossus.
- A second benchmark across 2,100 agent nodes exceeded 1 terabit per second of aggregate scan throughput. The comparable competitor test (8 read replicas) achieved less than 2× throughput while dropping primary throughput by more than 75%.
- Agent nodes access the full PostgreSQL engine: SQL, vector, full-text, spatial, columnar scans, and lakehouse federation, not a stripped-down or RAG-only interface.
- Agents connect through MCP; billing is per second of agent node activity; nodes auto-scale to zero when agents finish.
- The architecture depends on Google’s Colossus storage and Jupiter network; performance characteristics outside this vertically integrated stack are unknown. Agent nodes are currently read-only. The feature is in preview.
Work With Origins AI
Origins AI builds production AI systems for engineering teams. If your agents need low-latency access to live operational data without risking production stability, talk to our team.

