Detecting AI Agent Sandbox Escape: Agent Telemetry and Egress Alerts
AI agents escaping containment or misusing access produce telemetry you can watch. These detections focus on agent identities, network egress and shared infrastructure.
Signals worth watching
- Outbound connections from agent environments to destinations outside the allowlist.
- Agent workloads making requests through package proxies to unusual upstreams or non-package endpoints.
- Unexpected writes to shared registries, storage or repositories by agent identities.
- Agent identities attempting privilege changes, credential creation or access to new resources.
- Spikes in agent network actions, tool calls or compute usage.
- New communication patterns between agent runs.
Where the data lives
- Firewall and DNS logs (Azure Firewall, AWS Network Firewall, Route 53 Resolver, Entra Internet Access).
- Package proxy and registry logs.
- Cloud audit logs for agent identities (Entra sign-in and audit logs, CloudTrail).
- Agent platform logs (tool invocations, prompts and actions).
A starting query
Azure Firewall denials from agent subnets (structured logs):
AZFWApplicationRule
| where SourceIp startswith "10.50."
| where Action == "Deny"
| summarize Attempts = count(), Targets = make_set(Fqdn, 20) by SourceIp, bin(TimeGenerated, 15m)
Agent identities making directory or permission changes:
AuditLogs
| where InitiatedBy has "agent" or InitiatedBy has "svc-ai"
| where OperationName has_any ("Add", "Update", "Credential", "role")
| project TimeGenerated, OperationName, InitiatedBy, TargetResources
Adjust subnet ranges and naming to your environment.
Response
- Trigger the kill switch: revoke agent identity, block network, stop compute.
- Preserve logs and agent transcripts.
- Determine what was accessed outside the sandbox.
- Notify affected third parties promptly — delays compound the damage.