Google Extends Private AI Compute With Encrypted Server-Side Memory That Keys Stay Off Google’s Servers
On September 23, 2026, Google DeepMind published a technical update describing how Private AI Compute will gain persistent server-side memory. The announcement resolves a constraint that has shaped every prior design of cloud-based personal AI: choosing between remembering you across sessions or keeping your data private from the cloud operator. Google’s architecture attempts both. The cryptographic keys required to read your memory never leave your personal devices. The processing happens inside hardware-isolated cloud enclaves that Google itself cannot inspect during execution. The memory persists across sessions and across devices.
The Dilemma Private AI Has Always Faced
Local, on-device AI processing has been the cleanest privacy model available. What never leaves the device cannot be read by anyone else. The trade-off is compute: frontier AI models require substantially more processing than a phone or laptop can supply. The typical solution has been to offload computation to cloud servers, which means the cloud operator can, in principle, read whatever the model receives.
Google’s existing Private AI Compute platform addressed part of this by running cloud inference inside hardware-isolated enclaves, dedicated execution environments that are isolated from the host operating system and from Google’s own infrastructure. In that design, inference happens in a protected space, and the data processed during a session is never exposed to Google’s servers in a readable form during execution. However, the existing implementation was strictly stateless. The moment a session ended, all context was wiped. The enclave did not retain anything about you for future sessions.
That statelessness is a real limitation for personal AI. Knowing that you prefer certain approaches, that you are mid-way through a project, that you viewed assembly instructions on your smart glasses two days ago, these require memory that persists across sessions and across devices. Workarounds like having the model remember a list of facts about you exist, but they are shallow approximations of genuine continuity. The new architecture changes this by adding a persistent memory layer while keeping the same privacy constraints as local processing.
How the Architecture Works
The design introduces what Google describes as “a secure digital vault in the cloud.” Personal memory is stored in dedicated, encrypted cloud storage. The cryptographic keys required to decrypt that storage are derived from and held exclusively on the user’s personal devices. Google cannot access the keys, and therefore cannot read the stored memory without cooperation from the user’s device.
When an AI model needs to access memory to handle a request, the user’s device establishes an authenticated, end-to-end encrypted channel to a protected, isolated environment in the cloud. That environment, a hardware-isolated secure enclave, temporarily decrypts the memory in isolated memory space, processes the request, saves any new context back to encrypted storage, and immediately re-encrypts. The decrypted data exists only during the brief period the enclave processes the request, and it exists only inside memory that the enclave’s hardware guarantees cannot be accessed from outside.
Google’s architecture description references specific cryptographic components visible in their published diagram: a Key Encryption Key (KEK) and a Data Encryption Key (DEK), with a Wrapped DEK flowing through the system. The KEK, held on the user’s device, is used to encrypt the DEK. The encrypted or “wrapped” DEK travels with the memory storage. The enclave, once authenticated by the device, can temporarily unwrap the DEK and use it to access memory. After the operation, the DEK is re-encrypted and the memory is sealed again. Neither the KEK nor the plaintext DEK ever exist on servers outside the authenticated enclave during operation.
The design co-developed by Google DeepMind, Platforms and Devices, Core and Cloud teams, with executive sponsorship from Four Flynn, Jay Yagnik, and David Kleidermacher, applies the same hardware-isolation guarantee that previously made stateless inference private to a new problem: stateful, persistent memory that follows the user across their devices.
Trust Infrastructure: Transparency and Verification
The privacy model depends on trusting that the enclave is running the software Google claims it runs. Google addresses this through a tamper-proof public record of server software. User devices will verify the authenticity and integrity of the server-side software against this public record before transmitting any personal data. If the software on the server does not match the published record, the device does not establish the encrypted channel.
Google is publishing an updated Private AI Compute Technical Brief that covers system architecture, security proofs, and verification protocols, along with results from an independent audit by a cybersecurity firm that Google did not name in the announcement. The brief describes the full trust chain: hardware-enforced isolation, per-user databases protected by device-derived keys, authenticated channels between device and enclave, and the public software attestation log.
This combination separates three roles that are frequently conflated in cloud services: the infrastructure operator (Google), who owns and runs the servers; the enclave, which processes data during a request; and the user, who holds the keys. The infrastructure operator can see encrypted ciphertext and can measure that an enclave with a specific software hash ran. The operator cannot read what the enclave processed, because the keys never reside on infrastructure the operator controls outside the authenticated enclave. The user controls access by controlling the devices that hold the KEK.
Limitations and Open Questions
This announcement describes architecture and intent rather than deployed capability at scale. Google says the persistent memory feature will be added to Private AI Compute; the announcement does not identify specific products that have shipped it, a rollout timeline, or evidence from extended real-world use. The gap between a sound architectural description and a working production system is real, and the trust model assumes correct implementation of each layer of the stack.
The privacy guarantees depend on the correctness and security of multiple components: the enclave hardware (specific to the CPU manufacturer and the specific hardware generation), the key management implementation, the attestation protocol, the client software on user devices, and the server software running inside the enclave. A vulnerability in any layer is a potential privacy failure. The independent security audit is a meaningful step toward verifying the implementation, but the audit scope and findings are not yet public.
The architecture also raises a practical question: what happens when a user loses all their personal devices simultaneously? If the KEK is device-derived and held exclusively on those devices, a total device loss could mean permanent loss of access to the cloud memory. Google’s announcement does not describe recovery mechanisms, which suggests either that recovery paths exist but were not described in this post, or that the design accepts this trade-off in exchange for strict key custody.
For researchers interested in independent related work, the arXiv paper Opal: Private Memory for Personal AI (April 2026) presents a research design for personal AI memory using oblivious memory access inside trusted enclaves, which addresses some related threat models. It is not a Google system but describes the broader research landscape around this problem.
What This Means for Engineering Teams
The most direct implication is structural: Google is building a topology for personal AI that sits between device-only processing and ordinary cloud processing. Teams designing personal AI products now have a reference architecture that suggests how to separate user-owned cryptographic state (device-held keys) from cloud-stored encrypted data from inference compute. Whether to adopt Private AI Compute as a platform or to build similar trust topologies independently, the architecture is a concrete example of how the three elements can compose.
For teams building on AI infrastructure, the key architectural signal is the separation of model weights, temporary context, and durable user memory as three distinct components with distinct trust requirements. In the current standard cloud architecture, all three reside in the cloud under the operator’s control. Google’s design treats model weights as infrastructure, temporary context as ephemeral to the enclave, and durable user memory as user-owned ciphertext. That decomposition is available to any team willing to build the key management and enclave infrastructure, Private AI Compute makes it a platform rather than a from-scratch engineering project.
For teams evaluating AI model deployments that require user data privacy, this architecture demonstrates that the stateless-versus-stateful dilemma in privacy-preserving inference is tractable with current hardware. The challenge has moved from “is this theoretically possible” to “can we implement and audit it correctly,” which is a significantly more approachable engineering problem.
Key Takeaways
- Google’s Private AI Compute was previously stateless: context was wiped after every session. The new persistent memory layer adds cross-session, cross-device memory while maintaining the same privacy guarantees as local device processing.
- Encryption keys (KEK) are held exclusively on user devices; a wrapped DEK travels with encrypted cloud storage; the plaintext DEK exists only during brief decryption inside the hardware-isolated enclave.
- User devices verify server software against a tamper-proof public attestation log before sending any personal data, preventing the use of unauthorized or modified server software.
- An independent cybersecurity firm conducted an audit; updated technical documentation including architecture, security proofs, and verification protocols has been published.
- This is an architectural capability announcement, not a measured production deployment; privacy guarantees depend on correct implementation across hardware, key management, attestation, and both client and server software layers.
- Google describes cross-device use cases: resuming work seen on smart glasses from a laptop, or continuing conversations between mobile and web, scenarios that require persistent memory without exposing that memory to the operator.
Work With Origins AI
Origins AI builds production AI systems for engineering teams. If you are designing personal AI assistants that require strong user data privacy guarantees across devices, talk to our team.

