An AWS bill rarely grows because of one big decision. It grows through dozens of small ones: an instance sized for a launch that never came, a test environment nobody turned off, logs kept forever by default, traffic quietly routed through a NAT Gateway. The good news is that most of these are easy to find and safe to fix. In the accounts we review, the first pass often removes a meaningful share of the bill without touching application performance.
Start with Visibility
- Open Cost Explorer and group costs by service, then by usage type. The top three or four lines usually explain most of the bill.
- Tag resources with at least environment, project, and owner, and activate those tags as cost allocation tags so they appear in reports.
- Create AWS Budgets with alerts at 80 and 100 percent of the expected monthly spend, sent to people who can act on them.
- Turn on Cost Anomaly Detection so a sudden spike is reported within a day instead of at the end of the month.
- Review the bill monthly with engineering and finance together, so cost becomes part of technical decisions.
Right-Size Before You Commit
Many instances and databases are sized for a peak that rarely happens. Look at CPU and memory over at least two weeks, including busy periods, and use AWS Compute Optimizer for recommendations. An instance that averages 10 percent CPU can usually drop one size. Memory is not visible in default EC2 metrics, so install the CloudWatch agent before deciding. Do this before buying any commitment, otherwise you lock in a discount on capacity you do not need.
Pay Less for Steady Workloads
| Option | Discount | Use it for |
|---|---|---|
| Compute Savings Plans | Up to 66 percent versus On-Demand | Steady usage that may move between instance types, regions, Fargate, or Lambda |
| EC2 Instance Savings Plans | Up to 72 percent versus On-Demand | Steady usage that stays in one instance family and region |
| Spot Instances | Up to 90 percent versus On-Demand | Interruptible work such as CI runners, batch jobs, and stateless workers behind a queue |
| Graviton (Arm) instances | Lower price and often better performance per instance | Most Linux workloads in interpreted or JVM languages and Go, after testing |
Commit to the baseline you are confident about for the next one to three years, not to the current total. Leave the variable part on On-Demand or Spot. Spot instances can be reclaimed with two minutes of notice, so they suit work that can be retried, not a single database server.
Storage Savings That Are Easy to Miss
- Move EBS volumes from gp2 to gp3. The price per GB is about 20 percent lower, and the change can be made without downtime.
- Delete unattached EBS volumes and old snapshots. They are left behind every time an instance is terminated without care.
- Add S3 lifecycle rules to move old objects to cheaper storage classes, or use Intelligent-Tiering when access patterns are unpredictable.
- Add a lifecycle rule that aborts incomplete multipart uploads, which otherwise keep charging for invisible data.
- Set retention on CloudWatch Logs groups. The default is to keep logs forever, and log storage grows quietly every month.
Network Costs: NAT Gateway and Data Transfer
Network charges are the most common surprise. A NAT Gateway charges per hour and per GB processed, so private instances that pull container images, call S3, or download updates through it can generate a large bill. Add VPC gateway endpoints for S3 and DynamoDB, which have no charge and keep that traffic off the NAT Gateway. Serve static files and downloads through CloudFront, and keep services that talk to each other heavily in the same Availability Zone where reliability allows. Since 2024, every public IPv4 address is also charged by the hour, so remove unused Elastic IPs and avoid giving public addresses to instances that do not need them.
Turn Off What Nobody Uses
Development and staging environments are usually used during working hours only. Running them 12 hours a day on weekdays instead of around the clock removes about 64 percent of their compute hours. AWS Instance Scheduler or a simple scheduled Lambda can stop EC2 and RDS instances in the evening and start them in the morning. Also look for load balancers with no targets, old test stacks, unused RDS instances, and forgotten resources in regions your team does not normally use.
A Practical Monthly Routine
- 1Check the cost trend and anomalies, and explain every service that grew more than expected.
- 2Review Compute Optimizer and Trusted Advisor recommendations, and act on the safest ones first.
- 3Clean up untagged resources, unattached volumes, old snapshots, and idle IP addresses.
- 4Check Savings Plans utilisation and coverage before buying more commitment.
- 5Record what was changed and how much it saved, so the next review starts from facts.
Make every change reversible and measure it for a week. Downsizing an instance or a database is safe when you watch latency and error rates afterwards and can scale back up in minutes.
Key takeaways
- Make costs visible first with tags, budgets, anomaly detection, and a monthly review.
- Right-size instances and databases before buying Savings Plans, then commit only to the steady baseline.
- Switch to gp3, clean up volumes and snapshots, and set retention on logs and S3 objects.
- Reduce NAT Gateway, data transfer, and public IPv4 charges with VPC endpoints, CloudFront, and cleanup.
- Schedule non-production environments to run only during working hours.


