
Ask an engineering leader how their team is doing with AI tools and you will almost never get a stage. You get a mood: "we're figuring it out," "pretty good, actually," "a bit behind, honestly." None of those tell you what to do next, because none of them describe a state you could move out of on purpose.
A maturity model fixes that by naming the stages. This post lays out five for AI literacy specifically, at the level of a whole engineering team rather than one candidate, gives you a way to tell which one you are actually in without a survey, and says what it takes to move up one.
The format is not new. Carnegie Mellon's Software Engineering Institute built the original Capability Maturity Model in the 1980s to grade software development processes from ad hoc to optimized, and it became the template for grading almost any organizational capability the same way: security, data governance, DevOps practice (CMMI Institute). The useful idea underneath all of it is simple: capability grows in a predictable sequence, and you cannot skip a stage by buying a tool.
AI literacy fits the shape well, because it fails in exactly the pattern a maturity model is built to catch: an organization can look advanced (heavy tool usage, a Copilot license for everyone) while being immature (no one checks output consistently, no two engineers verify the same way). Usage and literacy are different axes, and a maturity model is what forces you to look at both.
One thing this model is not: it is not the four-behavior rubric we use to score one engineer's AI practice (deciding when to use a model, feeding it context, checking its output, owning the result), which we cover in detail in our AI literacy framework and rubric. That rubric grades a person. This one grades a team, and it is built out of how consistently that per-person rubric would actually land if you ran it on everyone right now.
| Stage | What it looks like day to day | How hiring happens | How training happens |
|---|---|---|---|
| 1. Unmanaged | People use AI tools quietly, inconsistently, sometimes against policy nobody enforces. No shared vocabulary for what "good" use looks like. | AI is not mentioned, or mentioned and not evaluated. | None. Whatever people pick up on their own. |
| 2. Ad hoc | Tools are sanctioned and widely used. Enthusiasm is high. Verification habits vary wildly by individual, and nobody has measured that variance. | "Comfortable with AI tools" appears in job ads with no way to check it. | A one-off lunch-and-learn or a vendor demo. Feels like progress, changes little. |
| 3. Defined | A written expectation exists: what can be pasted where, what must be checked before merge. Most engineers can describe it. Not everyone follows it the same way. | Interviews ask about AI use, but scoring is subjective and interviewer-dependent. | A course or rubric exists, but nobody has verified it changes behavior. |
| 4. Measured | The team scores actual AI-assisted work, not self-reports, against a consistent rubric. Gaps between individuals are visible and named. | Candidates are assessed on a real, observed task using the same rubric the team is scored on. | Training targets the specific gap the measurement found, usually verification, not tool tips. |
| 5. Adaptive | The rubric itself gets revisited as models and workflows change. Verification practice is a normal part of code review, not a separate initiative. Leadership can state the team's current gap in one sentence. | Hiring bar and internal bar are the same bar, and both move together as the rubric is updated. | Training is continuous and targeted, not an event. |
Two things about this table are easy to miss on a first read. First, stage 2 is where most engineering teams that describe themselves as "ahead on AI" actually sit: high usage, real enthusiasm, zero consistent measurement. Second, the jump from stage 3 to stage 4 is the one almost nobody makes, because it requires watching real work rather than trusting a written policy or a self-assessment, and that is a genuinely different kind of exercise than writing the policy was.
Do not run a survey. A survey asks people to rate their own literacy, and self-rated AI competence correlates weakly with actual verification behavior, for the same reason self-rated driving ability skews high in every study that has ever asked. Instead, check three things you can look at this week.
Pull up your last three postmortems or incident reviews and search for AI in the timeline. If AI-assisted code contributed to an incident, does the writeup say what was verified and what was not, or does it just note "AI-generated code was involved" and move on? The first is stage 4 language. The second is stage 2, dressed up as an incident review.
Ask two engineers, separately, how they check AI output before merging. Not whether they check it, everyone will say yes. Ask specifically what they check for. If the answers are close to identical and specific (both mention checking that a called function actually exists, or running the real failing case, not just the happy path), you are at stage 3 or above. If the answers are both "I read it over" with nothing more specific, you are at stage 2 regardless of how confidently they say it.
Look at your last few hiring debriefs for an AI-related round. If the scoring notes read like "seems comfortable with AI, uses it daily," that is not evidence, it is a vibe with a checkbox next to it. If they read like "walked through what they kept and why, caught that the suggested retry logic swallowed the original error," you have a rubric that is actually being applied. AI Literacy vs. AI Fluency goes into why the fluent-sounding answer and the literate one are easy to mix up in exactly this kind of note.
If you had to guess a stage from those three checks and you are not sure, you are probably at stage 2. That is not an insult. It is the single most common resting place, precisely because it requires no new process to arrive at, only enough tool access and enough time.
The move from 1 to 2 is a policy plus visibility problem: name what is and is not allowed, and make it known, which mostly closes the gap between quiet unsanctioned use and open sanctioned use.
The move from 2 to 3 is a definition problem: write down, in one page, what checking output actually means for your stack (which errors are common, which classes of bug generated code tends to hide), rather than leaving "review the code" to mean something different to every engineer. This is exactly the gap our AI literacy rubric is built to close, because a definition that everyone reads the same way is what a rubric actually is.
The move from 3 to 4 is a measurement problem, and it is the hardest one, because it means replacing "we have a written expectation" with "we checked whether people follow it." That requires watching real work rather than trusting a self-report or a policy doc, on a live task with the same tools the person would normally reach for. A 2025 randomized study from METR found that experienced open-source developers using AI assistance on real tasks took 19% longer than they would have unassisted, while believing they had been faster. That gap between belief and outcome is exactly what self-report cannot catch and observed work can. It is also the reason the EU AI Act's Article 4 obligation, which requires organizations deploying AI systems to ensure staff have "a sufficient level of AI literacy," has applied since February 2025: a training record is not evidence of literacy, and regulators are increasingly the ones asking for the difference.
The move from 4 to 5 is mostly discipline: keep re-running the same measurement as models change, because a rubric written for last year's failure modes quietly stops measuring anything once the tools improve past it.
Most of the work in stages 1 through 3 is writing things down. Stage 4 is different: it needs a way to watch someone do the work, not describe it, whether that person is a current engineer or a candidate. That is the part EasyEnv is built around. Candidates and engineers get a real machine with a real, broken task, tools including AI allowed, session recorded end to end. Scoring the four behaviors, deciding when to use the model, feeding it context, checking its answer, owning the result, becomes something you did by watching a recording, not something you inferred from a conversation. How to test if an engineer knows how to use AI walks through what that session actually looks like.
Being honest about the limit: this only pays off once you have something at stage 3 to measure against, a definition of what good checking looks like. Point a recorded assessment at a team with no shared definition yet and you will get interesting recordings and no way to score them consistently. Write the one-page definition first. The measurement is the next step, not a replacement for it.
Nobody needs to reach stage 5 to get the value out of this model. Stage 4, where you are actually measuring real work against a shared definition instead of trusting a policy or a self-report, is where the payoff is: it is the first stage where you can tell a training budget is working, tell a hiring bar is consistent across interviewers, and catch the gap between "we use AI a lot" and "we verify AI a lot" before it shows up in a postmortem instead.
So pick one of the three checks above and run it this week, on your own team, before you decide which stage you are in from memory.
Run live coding sessions and take-home challenges in real production environments. Watch sessions back, score consistently, and hire with confidence.
More posts you might like
Prompt engineering did not disappear when models got better, it changed shape. Here is what the skill actually looks like in 2026, why a resume line or a quiz cannot measure it, and the tasks and rubric that do.
CoderPad is a strong, well-built live coding tool for teams with a steady interview pipeline and a recruiting function to run it. A ten-person startup doing its own hiring is a different buyer with different constraints. Here is what changes, and what to check in any CoderPad alternative for startups before you commit.
Read moreEvery vendor calls its test a practical developer assessment and none of them define the word. Here is a four-question test you can run on any assessment in five minutes, worked through on a real example, to tell a work sample from a puzzle wearing a nicer outfit.
Read more