How to Use Amazon Macie to Discover PII in Your S3 Buckets
Retrospective: this article looks back at events from August 2017, written in 2026 with the benefit of hindsight.
Amazon Macie scans S3 buckets for sensitive data such as names, financial information and credentials. Here is how to run it effectively without a surprise bill.
Step 1: Enable Macie centrally
In AWS Organizations, designate a security tooling account as the Macie delegated administrator and enable Macie for member accounts. Findings then roll up to one place.
Step 2: Review the free bucket inventory
Once enabled, Macie automatically evaluates bucket security and access settings — which buckets are public, shared, unencrypted — before you scan any objects. Fix those findings first.
Step 3: Use automated sensitive data discovery
Turn on automated sensitive data discovery, which samples objects across your buckets and builds a sensitivity score per bucket. It is designed to give broad coverage at controlled cost.
Step 4: Run targeted discovery jobs
For buckets that score high or are known to hold customer data, create a sensitive data discovery job:
- Scope to specific buckets and prefixes.
- Use managed data identifiers for common PII and financial data.
- Add custom data identifiers for your own formats, such as customer or policy numbers.
- Exclude file types you do not need to scan, such as images or compressed logs.
Step 5: Act on findings
Send Macie findings to Security Hub and route high-severity findings to a ticket queue. For each sensitive bucket: confirm the owner, restrict access, enforce encryption and decide on retention.
Common mistakes
- Scanning everything on day one and generating a large bill.
- Producing findings nobody owns. Assign data owners before you scan.
- Amazon Macie Launches (Aug 2017): Machine Learning for Finding Sensitive Data in S3 Platform Changes
- Amazon Macie Rollout Checklist: Cost Controls and Finding Triage How-To & Hardening
- CIO Brief: You Can't Protect Data You Haven't Found CIO Briefings