How to reduce AWS costs without slowing your team down
When the AWS bill grows faster than your traffic, the cause is rarely one big mistake. It’s dozens of reasonable defaults that nobody revisited. Here is where cloud spend usually hides, how engineering teams find it, and which cuts are safe.
Written for CTOs, founders, and engineering leads who own an AWS bill and need it to make sense.
- ComputeEC2 · ECS / Fargate · Lambda · EKS
- Instances sized for a peak that never comes
- Dev and staging running 24/7
- Old instance generations
- No commitment on a stable baseline
- DatabaseRDS · Aurora · DynamoDB · ElastiCache
- Oversized instances hiding slow queries
- Provisioned IOPS nobody needs
- Idle replicas and forgotten clusters
- Data transferNAT Gateway · cross-AZ · egress
- S3 and API traffic routed through NAT
- Chatty services across availability zones
- Uncached downloads
- StorageS3 · EBS · snapshots
- No lifecycle rules
- Unattached volumes
- Years of snapshots
- ObservabilityCloudWatch Logs · metrics
- Debug logging in production
- Retention set to “never expire”
- High-cardinality custom metrics
- Idle resourcesLoad balancers · IPs · test stacks
- Load balancers with no targets
- Public IPv4 addresses (charged since 2024)
- Abandoned experiments
Tile size suggests where teams usually look first, not measured shares of any bill.
Symptoms of an AWS bill that has drifted
- Cost grows faster than usageTraffic is up modestly, but the bill keeps climbing month after month.
- Nobody can explain the billYou can’t say which product, team, or customer drives which part of the spend.
- Non-production costs as much as productionStaging, QA, and developer environments run around the clock at production size.
- Surprise line itemsNAT Gateway, data transfer, or CloudWatch charges appear larger than the servers.
- Commitments that don’t fitReserved capacity sits unused, or there’s none at all on a stable workload.
- Cost only reviewed at renewal timeThe bill gets attention quarterly, after the money is already spent.
Why AWS costs drift
Cloud cost is an engineering property, like performance or security. It drifts for predictable reasons:
- Sized for the worst day. Capacity chosen during a launch or an incident is rarely scaled back afterward.
- No ownership. Without consistent tagging, spend belongs to everyone and therefore to no one.
- Architecture side effects. Private subnets route S3 traffic through a NAT gateway billed per gigabyte. Services talk across availability zones, and every extra microservice brings its own compute, load balancer, and logs. Verbose logs are ingested and kept forever. Each choice was reasonable; together they add up.
- Data only grows. Objects, snapshots, and logs accumulate without lifecycle rules.
- Pricing changes. AWS pricing evolves, for example the per-address charge for public IPv4 introduced in 2024. Setups that were cheap when designed may not be cheap now.
The business risk isn’t just the bill
Unexplained cloud spend erodes margin quietly, but the bigger risk is what happens when it finally gets noticed. Rushed cuts, such as turning off redundancy, shrinking databases before a busy season, or deleting backups, trade a cost problem for a reliability problem. Or the savings come out of the product roadmap instead of the infrastructure.
A deliberate approach separates waste (spend that buys nothing) from headroom (spend that buys resilience), and removes the first without touching the second.
Where engineering teams usually find unnecessary cost
The six areas on the cost map, in detail: what to look for, how to find it, and the risk of fixing it.
| Area | Common waste | How to find it | Typical fix | Risk of the fix |
|---|---|---|---|---|
| Compute | Low average utilization; always-on non-production | CloudWatch utilization, AWS Compute Optimizer | Rightsize, schedule non-production, newer or Graviton instance types | Low to medium: test under load first |
| Database | Oversized instances covering for slow queries | Performance Insights, slow-query logs, EXPLAIN plans | Fix queries and indexes, then downsize; review provisioned IOPS | Medium: change during quiet hours with a rollback |
| Data transfer | NAT processing and cross-AZ charges | Cost Explorer by usage type, VPC Flow Logs | S3 gateway endpoints, zone-aware routing, caching and CDN | Low to medium |
| Storage | Old snapshots, unattached volumes, cold data in hot tiers | Cost Explorer, S3 Storage Lens, EBS inventory | Lifecycle rules, Intelligent-Tiering, gp2 to gp3, delete orphans | Low, once retention needs are confirmed |
| Observability | Verbose log ingestion, infinite retention | CloudWatch usage by log group | Lower log levels, set retention, sample high-volume logs | Low: keep what incidents actually need |
| Idle resources | Unused load balancers, IPs, test stacks | Trusted Advisor, tagging reports | Delete, and add expiry tags to experiments | Low |
A safe order of operations
Visibility first, commitments last. Committing before cleaning up locks in today’s waste.
- Visibility
Enforce cost allocation tags in infrastructure code; group Cost Explorer by service and tag.
- Remove waste
Idle, orphaned, and forgotten resources. Low risk, immediate effect.
- Rightsize
Match instance and database sizes to measured load, with headroom.
- Fix architecture
Data transfer paths, caching, storage tiers, and logging.
- Commit
Savings Plans or Reserved Instances for the stable baseline.
- Guardrails
Budgets, anomaly detection, and cost review in normal engineering work.
Savings options compared
Each lever saves money in a different way, and each has a cost of its own.
| Option | How it saves | Commitment | Best for | Tradeoff |
|---|---|---|---|---|
| Rightsizing | Pay for capacity you use | None | Anything sized by guesswork | Needs load data and testing |
| Scheduling | Stop resources outside working hours | None | Development, QA, and staging | Environments are not instantly available at night |
| Savings Plans | Discount in exchange for committed spend | 1 or 3 years | Stable compute baselines | Over-committing wastes money |
| Reserved Instances | Discount on specific capacity | 1 or 3 years | Steady databases and caches | Less flexible than Savings Plans |
| Spot capacity | Use spare capacity at a steep discount | None | Interruptible batch and CI workloads | Instances can be reclaimed with short notice |
| Graviton (ARM) | Often better price-performance | None | Workloads that build for ARM | Needs compatible images and testing |
Discount levels vary by service, region, and term. Check current AWS pricing rather than relying on rules of thumb.
What’s safe to cut, and what to cut deliberately
Cut deliberately, with data
- Multi-AZ redundancy for production databases
- Backup retention and disaster-recovery copies
- Production capacity headroom before busy periods
- Logs and metrics your incident response depends on
Usually safe to cut
- Idle load balancers, IPs, and abandoned stacks
- Unattached volumes and snapshots past retention
- Non-production running outside working hours
- Debug-level logging in production
Quick AWS cost checklist
- Cost allocation tags enforced in Terraform or CloudFormation
- AWS Budgets and Cost Anomaly Detection alerts on
- S3 gateway endpoint in every VPC that uses S3
- Log group retention set everywhere
- gp3 volumes instead of gp2
- Non-production environments scheduled
- Compute Optimizer recommendations reviewed
- Commitments sized to the post-cleanup baseline
How Yippify can help
Yippify reviews AWS accounts the way engineers do: from the bill down to the architecture decisions behind it. We identify waste, rank the fixes by savings against risk, implement them in infrastructure code with your team, and put guardrails in place so the bill stays explainable. If the answer is that the spend is justified, we’ll say so.
Sometimes the bill is a symptom of an aging architecture rather than idle resources. Then the real decision is whether to modernize the system or replace it, which we cover in should you rewrite your application from scratch?
Questions teams ask
Where does most AWS waste come from?
Usually from resources sized for peak load and never revisited, environments that run around the clock when they are only used during work hours, storage that grows without lifecycle rules, data transfer through NAT gateways or across availability zones, and logging retained or ingested at a higher level than anyone reads.
Should we buy Savings Plans or Reserved Instances first?
Rightsize and remove waste first, then commit. Committing to today’s usage locks in today’s inefficiency. Once usage is stable, Savings Plans or Reserved Instances covering the predictable baseline are usually one of the larger single levers.
Will cutting costs hurt reliability?
It should not if you separate waste from headroom. Removing idle resources, lifecycle rules, and right-sized development environments carry little risk. Reducing production capacity or redundancy is a deliberate tradeoff that should be made with load data and a rollback plan.
How do we find out which team or feature is driving the bill?
Consistent cost allocation tags enforced in infrastructure code, AWS Cost Explorer grouped by tag and service, and the Cost and Usage Report data exported for deeper analysis. Without tagging, most cost conversations become guesses.
Is moving off AWS the answer to high costs?
Occasionally, for steady workloads at scale, but it trades cloud spend for hardware, staffing, and operational risk. For most teams, fixing architecture and usage patterns within AWS is faster and lower-risk.
How quickly can costs come down?
Removing clearly idle resources and fixing obvious configuration can take effect within days. Architectural changes, such as caching, reducing cross-zone traffic, or re-platforming a workload, take longer and should be prioritized by how much they save against how much risk they add.
AWS bill growing faster than your traffic? Let’s look at where the cost is coming from.
Tell us roughly what you spend each month, which services dominate the bill, and whether anything changed recently. We’ll point to the likeliest places to look first.
- Waste separated from headroom
- Fixes ranked by savings and risk
- Changes made in infrastructure code
A rough description is enough to start. No specification needed.