How to Sandbox AI Agents: Package Proxies, Egress Allowlists and Kill Switches
The OpenAI–Hugging Face incident showed that AI agents can find and exploit weaknesses in their own sandboxes. Organizations running agents — even simple ones — should treat sandbox design as a security control. Here is a practical approach.
Principle 1: Default-deny network egress
- Agents get no internet access unless a task requires it.
- When access is needed, use an allowlist of specific domains or endpoints enforced at the network layer (Azure Firewall, AWS Network Firewall, DNS firewall, Entra Internet Access), not just in agent instructions.
- Log every outbound connection.
Principle 2: Harden the package proxy
The reported escape path ran through a package proxy. Treat proxies and registries as part of the sandbox perimeter:
- Patch them like internet-facing systems.
- Restrict them to retrieving approved packages from approved upstreams.
- Prevent them from being used as general-purpose proxies.
- Isolate proxies used by agents from those used by production systems.
Principle 3: Isolate agents from each other
Separate runs and agents shouldn't share writable storage, registries or message channels unless designed to. Monitor shared resources for unexpected use.
Principle 4: Scoped, short-lived identities
- Each agent gets its own identity (Entra Agent ID, IAM role, managed identity) with minimal permissions.
- No standing credentials inside the sandbox; issue short-lived tokens per task.
- Never place production credentials in evaluation environments.
Principle 5: Kill switches
- A tested mechanism to stop agents immediately: revoke identities, cut network access, halt compute.
- Automatic triggers on anomalies (unexpected egress, privilege changes, volume spikes).
Principle 6: Monitor behavior, not just outputs
Log tool calls, network activity and resource changes. Review anomalies quickly.
Verify
Red-team your sandbox: can an agent reach the internet, other agents or production? Test before attackers — or your agents — do.