Overview
- Google DeepMind announced the pilot on Thursday, August 27, 2026, and is running a Gemini Flash Lite model inside Google Cloud’s Confidential Space so evaluators never see model weights and Google never sees the evaluator’s test prompts.
- The evaluation runs inside a cryptographic enclave that stores model weights in hardware‑encrypted GPU memory and keeps prompts in encrypted host memory, with remote attestation used to verify the exact software environment.
- DeepMind is working with outside organizations including the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons to test private benchmarks, and the company says it will publish a technical report describing methods and findings.
- The pilot limits exposure by approving evaluation code, blocking external network access, and running on a single H100 confidential GPU for now, which highlights scaling limits and the need for legal and code reviews between parties.
- If adopted more widely, the method could provide stronger technical guarantees against benchmark leakage, protect intellectual property and sensitive test data, and make independent model scores more trustworthy for policymakers and enterprises.