Cloud 11 min read

Practical Cloud Cost Optimization That Survives Finance Reviews

Cloud bills grow quietly. This note orders cost work by payoff so you can cut waste this week, then decide whether a full FinOps program is actually necessary.

Start with the bill, not the architecture

Before changing infrastructure, export last month's cost by service, team, and environment. Most teams find three to five line items that explain the majority of growth: unattached volumes, oversized compute, idle load balancers, and log retention that nobody scoped.

Architecture rewrites are expensive. Deleting an unused database replica is not. Sequence the work by dollars per hour of effort, not by elegance.

Kill idle, then rightsize

Idle resources are pure waste: volumes with no attachment, load balancers with no targets, development clusters that run all weekend. Automate a weekly report of zero-utilization resources and make deletion a normal review item.

Rightsizing is next. Use two weeks of CPU, memory, and network percentiles — not averages — and step instances down only when p95 stays healthy after the change.

If you cannot name the workload owner and the performance budget, you are not rightsizing — you are guessing with someone else's latency.
  • Target p95 CPU under ~60% after rightsizing to keep headroom for bursts.
  • Prefer vertical downsize before horizontal scale-out on stable workloads.
  • Document the performance budget so finance reviews do not reverse the change later.

Commitments only after the baseline is stable

Savings plans and reserved instances look attractive on a slide and painful when your architecture is still moving. Lock commitments only for the stable base load — typically 60–70% of a predictable production floor.

Track commitment utilization as a product metric. Underused commitments are a tax on future change.

Make cost visible in the same places as reliability

Cost dashboards that live only in finance tools get ignored by the people who provision. Put a simple cost trend next to latency and error rate in the team's weekly review.

Tagging is not optional. Without owner, environment, and service tags, every optimization becomes an archaeology project.

  • Require tags at provisioning time; reject untagged resources in CI where possible.
  • Alert on week-over-week spend jumps for a service, not just monthly totals.
  • Give each squad a budget envelope and make overruns a planning conversation.

Key takeaways

  • Optimize by dollars-per-hour of effort: idle cleanup before redesign.
  • Rightsize from percentiles with a documented performance budget.
  • Commit only the stable base load and track utilization.
  • Surface cost next to reliability metrics or nobody owns it.
  • Tagging is the prerequisite for every other cost practice.