
Teams interview platform engineers on Kubernetes depth, Terraform fluency and networking, then act surprised when the monthly invoice climbs for two quarters and nobody says anything.
Cost awareness is an engineering skill. It is not finance, it is not a dashboard, and it is not somebody else's job. It is the same skill as performance work pointed at a different number, it is testable in twenty minutes on a running environment, and it is almost never tested.
It sounds like a finance topic. The word that gets used is FinOps, which sounds like a department, so it ends up owned by nobody in the engineering loop.
The knowledge part is trivially available. Reserved instances are cheaper than on demand. Storage tiers exist. Right size your requests. Every candidate can recite this and a model can recite it better, so interviewers correctly conclude the question is worthless and then drop the topic entirely rather than changing the question.
The consequences are slow. A bad latency decision shows up in an hour. A bad cost decision shows up in a quarter, by which time nobody connects it to the person who made it.
The result is a generation of infrastructure engineers who can build anything and have never once been asked what it costs.
Not knowing prices. Three things:
Noticing. Seeing that something is disproportionate without being told to look. A node pool at eight percent utilisation, a log pipeline ingesting debug output from production, a dev environment that has been running since March, a snapshot schedule nobody cancelled.
Attributing. Being able to say what is costing the money. In most organisations this is genuinely hard, which is why so many cost programmes start and die at "the bill went up".
Trading off correctly. This is the one that separates people, and it is where cost work goes wrong. Every saving is a reduction in something: headroom, redundancy, retention, latency, or somebody's time. An engineer who cuts spend without knowing which one they traded has not saved money, they have moved a risk somewhere less visible.
Give the candidate a running environment with real resource definitions and a handful of deliberate inefficiencies. Ask a plain question: what would you change here, what would it save, and what would it risk.
Inefficiencies that work well, mixed so not all are safe to cut:
Wildly oversized requests on a stateless service. Requests set at four times observed usage, so the scheduler reserves capacity nobody uses. Safe to cut, and the candidate should still ask about peak behaviour first.
A node pool sized for a workload that moved. Idle capacity that costs money every hour. Safe, and the interesting part is whether they check what else is scheduled there.
A log pipeline shipping debug output. Often the single largest line in a real bill and frequently invisible. Strong candidates ask what the retention requirement is before touching it, because sometimes the answer is a compliance obligation.
A latency sensitive service with tight requests. The trap. It looks over-provisioned by the same crude measure as the first item, and cutting it produces an incident under load. A candidate who treats all four the same way has told you what you needed to know.
Old snapshots and orphaned volumes. Cheap to fix, and a good test of whether they check what is attached before deleting anything.
Do they ask what the workload is before touching it? This is the single best predictor. Strong candidates want to know the traffic pattern, whether it is bursty, what the latency requirement is, and what happens at peak. Only then do they propose a change. Candidates who start trimming requests off a list of utilisation percentages are doing arithmetic, not engineering.
Do they look at utilisation over time or at a point? A snapshot of CPU at 2pm on a Tuesday says almost nothing about a service with a nightly batch job.
Do they quantify? Not to the cent. "This pool is about a third of the compute line and it is running at eight percent" is the level of precision that makes a decision possible. An engineer who cannot connect a resource to a rough share of the bill will never prioritise correctly.
Do they think about who pays the cost of their change? Reducing an environment from three replicas to two saves money and transfers risk to whoever is on call. Reducing retention from ninety days to seven saves money and transfers cost to whoever next debugs a slow-burn incident. Naming the transfer is the mark of somebody who has done this before.
Do they mention the non-infrastructure levers? The largest savings in most organisations are not in resource sizing. They are in things running that should not be: environments left up, jobs that no longer feed anything, data retained because deleting it needed a decision. Our own view of that is in stop paying for idle servers and cloud cost optimization, and the closest assessment sibling is the cost review exercise in the Terraform and AWS skills assessment.
A cost question asked in conversation returns the recital: reserved instances, right sizing, spot where appropriate. Asked in front of a running cluster with real manifests and real metrics, it returns judgment, because the candidate has to pick which of five plausible things to touch.
In EasyEnv the candidate gets a real environment they can inspect and change, with the inefficiencies you planted. The session is recorded, so you can review what they looked at before they proposed anything. As with interviewing a platform engineer on a real Kubernetes cluster, the entire value is that the system is real enough to push back.
It is also a good AI test almost by accident. A model will produce a confident, generic list of cost optimisations for any stack you name. It cannot tell which of these five services is latency sensitive. Watch whether the candidate goes and finds out.
Noticing: did they find the waste without being pointed at it?
Judgment: did they treat the risky cut differently from the safe one?
Quantification: can they say roughly what it saves?
Honesty about tradeoffs: did they name what each change costs in reliability, retention or human time?
A candidate who finds three of five inefficiencies and correctly refuses to touch the fourth has outscored one who proposes cutting all five.
Do not ask for pricing from memory. Prices change, regions differ, and looking them up is the correct behaviour.
Do not frame it as a savings target. "Cut thirty percent" produces the reckless answer, and rewards it. Ask what they would change and why, then ask what it risks.
Do not treat it as a separate interview. It is one section of a platform or SRE loop, worth twenty minutes, sitting next to the reliability work in what a real DevOps interview should look like.
What is the largest line on your cloud bill, and when did an engineer last look at it on purpose?
Run live coding sessions and take-home challenges in real production environments. Watch sessions back, score consistently, and hire with confidence.
More posts you might like
Under the EU AI Act, AI used to screen or evaluate candidates sits in the high risk category. Here is what that actually asks of you as an employer, the questions to put to your assessment vendor, and how to structure a process you could defend.
When the model has shell access, the question is no longer whether a candidate can write the code. It is whether they can supervise something that writes it faster than they can read it. Here is how to test that on a real repository.
Read moreObservability interviews usually become tool interviews, which sort candidates by which vendor their last employer bought. Here is how to test the actual skill: narrowing down a failure that only happens to four percent of requests.
Read more