AI Sandboxes: Sandboxes Begin Using Six-Dimensional Evidence to Constrain Embodied and Cyber-Physical Deployment Claims
Submitted on June 16, 2026, the AI Sandboxes study proposes a threat model, taxonomy, and measurement framework covering digital AI, embodied autonomy, and cyber-physical systems, and instantiates it with three real sandbox cases.
Sandboxes provide both isolation and determine the scope of applicability of test conclusions. For systems that perceive, decide, execute, and network, any weak boundary can invalidate safety claims.
The framework formalizes sandbox boundaries and the weakest-link rule, categorizing evidence into six dimensions: fidelity, controllability, observability, isolation, reproducibility, and governance artifacts, and includes attacks on the assurance infrastructure itself within the cyber-physical threat model and full verification scope.
Regulation and procurement for robotics, AIoT, and high-risk agents will require reproducible sandbox evidence rather than just model testing; the test environment itself will become critical infrastructure subject to audit.
High-risk projects should first define the risks that the sandbox can and cannot cover, preserve six-dimensional evidence and versions, and then limit deployment scope to within the validated claims.
Cross-industry standards, public measurement tools, and incident benchmarks need to be developed, and how sandbox evidence maps to real-world deployment risk and continuous monitoring must be validated.