> ## Content Index
> Fetch the complete content index at: https://www.shiftharness.tech/llms.txt
> Use this file to discover other available public pages before exploring further.

# The AI Adoption Maturity Ladder: L0 → L4
- URL: https://www.shiftharness.tech/ai-adoption-maturity-ladder-l0-l4/
- Published: 2026-08-21T09:23:55.000Z
- Updated: 2026-08-21T10:40:14.000Z
- Description: Most AI programs stall around month nine because nobody can show the board what actually changed. A five-rung ladder that places an org by the artifacts it produces - PRs, ADRs, test plans, postmortems - and names the missing operating asset at each rung.
- Author: Sergii
- Tags: AI Adoption, AI Maturity Model, AI transformation, AI operating model, Delivery Teams, AI in Software Delivery, AI Enablement, AI Strategy, AI Measurement, #pillar

Most boards looking at AI dashboards are looking at the wrong one. They count Copilot seats, weekly active users, prompts run per developer per week. The numbers go up. The delivery metrics do not. The CEO eventually asks the obvious question. *If everyone is using the AI tools, why hasn't anything actually moved?* Nobody has a defensible answer. Tool usage frequency is not the right signal for maturity, and confusing the two is why most [AI transformation programs](https://www.shiftharness.tech/ai-operating-model/) stall around month nine.

> **The AI adoption maturity model** is a five-rung framework that measures how deeply AI is embedded in how work gets done, not how often people open the tools. The rungs run from L0 (Awareness) through L4 (Measured, governed, continuously improved). The most reliable single place to start is the artifacts the team produces rather than what the team self-reports. Artifacts alone will not settle attribution or outcome, which is why the placement method below pairs them with tooling evidence.

One scope note before the rungs. This ladder measures how deeply AI has changed delivery work. It is not a total enterprise AI capability maturity model: governance and security, value measurement, data readiness, operational resilience, and workforce capability are parallel dimensions, each carrying a minimum gate that applies at every rung rather than arriving at the top one. Report the delivery rung and the capability gaps separately, or a well-integrated team with weak governance reads as mature when it is not.

In AI rollouts across delivery orgs, adoption keeps landing in the same five-rung shape. The ladder is a practitioner model drawn from repeated observation, not a validated instrument with published inter-rater data, and it is worth using as the former rather than citing as the latter. Each rung is defined by what you can actually see in PRs, test plans, specs, postmortems, and the surrounding process. Not by what people say they are doing.

## The next asset at each rung

Most AI-investment conversations confuse two questions: "what tool to buy next?" and "what operating asset to build next?" Only the second one moves the org up a rung, and it is the second one that belongs in front of a board.

| Current rung | What this rung is                                                                   | What to build                                                                                                                                           | Next investment                       |
| ------------ | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| L0           | Awareness only; nothing has moved                                                   | Tool access, safety policy, basic enablement                                                                                                            | Approved usage baseline               |
| L1           | Basic tool usage; private and invisible in artifacts                                | Role playbooks, examples, refreshed training (AI basics, tools like Claude, Codex, Cursor), personal productivity                                       | Role-specific workflow change         |
| L2           | Role-based workflow usage                                                           | Role-based process automation, Definition of Done, review checklists, templates, shared prompt libraries                                                | Process integration                   |
| L3           | Codified delivery system; AI designed into the pipeline, with exclusions on purpose | Spec-driven development, codified SDLC (defined skills, agents, pipelines, improvement loop, quality gates on CI/CD), improved cycle time and lead time | Instrumentation and feedback loops    |
| L4           | Measured, governed, continuously improved                                           | Full enablement of a code factory. Decision rights, cost/quality tuning, cross-project learning                                                         | Sustained optimization; no rung above |

Every rung in the rest of this article comes back to this table. Place the org on the ladder, then read across to the next investment on that row and fund it. The trending tool usually belongs to a rung the org has not reached yet.

## L0 - Awareness: people know the tools exist; nothing has actually moved

At L0, AI exists as awareness, procurement, or scattered experimentation. There's no consistent approved usage pattern, no role guidance, and no observable change in delivery artifacts. The team has heard about Copilot, Claude, Cursor, ChatGPT; there may be a Slack channel with a few links. The work product looks identical to what was shipping a year ago.

The signal for L0 is artifact-level homogeneity. Pull a sample of recent PRs, test plans, design docs, retros, postmortems and lay them side by side with the same sample from twelve months earlier. If they look indistinguishable in structure, depth, and the categories of decision they capture, the team is no higher than L1 regardless of what the procurement spreadsheet says about license counts. Separating L0 from L1 takes one more question: is anyone using AI at all, consistently and with approval? If not, it is L0.

The common failure mode at L0 is confusing tool procurement with capability. A purchase order is not an operating model change. A vendor pilot is not adoption. Most orgs that report "we are doing AI adoption" because they bought licenses are sitting at L0 and have not noticed.

## L1 - Basic tool usage: individuals open the tools and paste outputs back into the work

At L1, individuals actively use AI, but the usage remains private, inconsistent, and invisible in team-level process artifacts. AI-shaped fragments start showing up in commits and tickets; the work happens in parallel to AI, not through it.

AI involvement is invisible at the team level: no record of which spec was AI-assisted, no review checklist asking whether AI was used and verified, no measurement separating AI-assisted from non-AI work. Ask the team where AI shows up in their delivery process and the answer is a shrug, or a story about one person.

The common failure mode at L1 is the AI adoption dashboard declaring success. Weekly active users rise, license utilization looks healthy, prompt counts compound. None of it tells you whether the operating model has changed, because at L1 it has not. The org is measuring procurement and calling it transformation. This is the rung where boards lose patience around month nine.

Per-role, L1 is depressingly consistent. A developer autocompletes a function and ships it without the test the AI also offered. A QA pastes in a user story, gets back test cases, and picks the two they would have written anyway. A PM summarizes a stakeholder transcript and uses the bullets verbatim. A BA accepts the first draft of acceptance criteria. A solution architect uses it once for a diagram and never returns. None of these are wrong. They are simply not adoption.

## L2 - Role-based workflow usage: AI is embedded in how a role does its specific work

L2 is the first rung where workflow change becomes observable. It's not yet evidence that delivery improved. The shift is not that more people are using the tool more often. The shift is that AI has moved from a sidebar that an individual opens occasionally to a component of how a specific role does its specific work. A PM at L2 does not "use AI for some things"; a PM at L2 uses AI inside spec drafting, scope challenge, and risk surfacing. Those activities are different than they were a year ago, and the difference is visible in the artifacts.

This rung is where real **AI capability progression** starts. It's also where most organizations get stuck. The reason is almost always the same: training was delivered as a generic "how to prompt ChatGPT" session, role playbooks were never written, and seniors quietly dropped the tools because the workflow gain was not obvious for the kind of work they actually do. Reaching L2 requires role-level redesign, not more enthusiasm and not better access.

Here is what L2 looks like across the five delivery roles, in compressed form. The full per-role rubric - including evidence depth and how to score an individual - belongs in the companion article.

| Role      | L2 artifact signal                                                                                         |
| --------- | ---------------------------------------------------------------------------------------------------------- |
| Developer | PRs show AI-assisted implementation reasoning, test scaffolding, and review of generated paths             |
| QA        | Test plans show AI-assisted edge-case coverage, traceability, and defect-pattern awareness                 |
| PM        | Stories are smaller, sharper, and include assumptions, risks, and acceptance criteria strengthened with AI |
| BA        | Requirements include ambiguity checks, source traceability, and earlier clarification loops                |
| SA        | ADRs show evaluated alternatives, rejected options, NFR trade-offs, and AI-assisted risk analysis          |

The unifying signal is that the same person ships more useful work and the AI involvement is traceable in the artifacts.

Observable is not the same as better, and this is where adoption reporting quietly overreaches. The evidence runs as a chain: exposure, changed behaviour, artifact quality and active use, delivery and quality outcomes, then cost and risk guardrails. Each link is a separate measurement, and a break anywhere means the chain proves less than it appears to. [DORA's 2025 findings](https://dora.dev/insights/balancing-ai-tensions/?ref=shiftharness.tech) are the sharpest illustration: the research reversed the prior year's result and now associates AI adoption with higher delivery throughput while still associating it with lower delivery stability. That matters here because throughput is the number most likely to be offered as proof that AI worked, and it is not what this ladder places a team on. Read the stability half in the same review as the speed half, or the improvement story rests on the half of the finding that flatters it.

The common failure mode at L2 is that the org tried to skip past it. Six months after a generic rollout, AI use collapses back to the L1 pattern and the conclusion is "the tools must not be ready." The tools are usually fine. The role-level redesign was never done.

## L3 - Integrated into the delivery process: the pipeline is designed around AI, with exclusions on purpose

L3 is where role-level redesign hardens into process-level redesign. The pipeline (planning, spec, design, build, test, release, postmortem) has been redesigned on the assumption that AI is in it, which is not the same as AI running in every step. A mature team deliberately excludes AI from steps where the risk does not justify it, and records that decision; an immature one automates everything and calls the coverage maturity. Handoffs, definitions of done, and quality gates are all different. The acceptance criteria checklist a developer sees on opening a PR carries items that did not exist a year ago: was AI used in implementation, is the AI-suggested test coverage documented, has the AI-assisted code been reviewed against the team's prompt library.

Removing AI from the team's tooling tomorrow tells you something real, but it does not tell you the rung. At L2, some individuals slow down and the seniors barely notice, because nothing in the process assumed AI was there. At L3, removal forces the team to redesign the SDLC back to a previous state. That degradation confirms the dependence the process artifacts already implied. The process artifacts are what place the team. Standards documents have been rewritten. Role playbooks reference AI as a default. Prompt libraries, eval suites, and AI-aware code review checklists are part of the standard toolkit, not personal experiments.

Run the fallback question alongside the placement rather than inside it. Can the team operate without AI on a path it has actually tested, and say what that costs in service level, cycle time, and quality? A team that has rehearsed it and can put a number on it has a managed dependency. A team that has never tested it has an unmanaged one, and that belongs in the report next to the rung, not in the rung. An integrated team with an untested fallback is still at L3\. It is an L3 carrying a single point of failure, and saying both things is more useful to a leadership team than averaging them into one number.

The visible evidence for L3 is the process-artifact checklist. Inspect the team's templates and standards documents and ask: has each one been rewritten to be AI-aware?

| Process artifact    | AI-aware change                                                             |
| ------------------- | --------------------------------------------------------------------------- |
| Definition of Done  | Defines verification expectations for AI-assisted work                      |
| PR template         | Captures material AI assistance, generated tests, and review responsibility |
| Test plan template  | Includes AI-generated edge cases, traceability, and coverage rationale      |
| Story template      | Includes assumptions, risks, and AI-assisted completeness checks            |
| ADR template        | Includes AI-assisted alternatives and rejected options                      |
| Postmortem template | Captures whether AI contributed to or could have prevented the issue        |
| Role playbooks      | Define how each role uses and verifies AI-assisted work                     |

If five or more of these have been rewritten in the last two quarters and the team can show you the diff, treat that as a screening flag for L3 rather than a threshold that settles it, and confirm it against whether the rewritten templates are actually in use. If two or fewer, the team is at L2 with L3 ambitions.

The evidence at L3 is process documentation, not individual artifacts. The team also typically has a small but real prompt library: a shared, versioned asset that lives in the same repository as the rest of the engineering standards.

This is the rung where **AI transformation maturity** stops being about individuals and starts being about the system. What defines it is a codified delivery system: spec-driven development, an SDLC written down rather than remembered, with defined skills, agents and pipelines, an improvement loop, and quality gates wired into CI/CD rather than carried in someone's head. Wired in is not the same as enforced: a check blocks only when the pipeline is configured to require its result and bypass authority is controlled, so record which gates are advisory and which actually stop a merge. Where it lands, cycle time and lead time are usually where it shows up first. The unglamorous infrastructure (playbooks, libraries, templates, checklists, training that gets refreshed rather than delivered once) is what makes the L2 changes reproducible by a new hire in their first month.

The worst version of L3 I keep seeing is a delivery system that gets codified and then never instrumented. The team rewrites the templates, publishes the prompt library, updates the definition of done, wires the quality gates, and then the data layer never gets touched. No telemetry tells the org whether the new process is producing the outcomes it was meant to produce. The work was done. The signal that would confirm it landed never got built. That is exactly the L3 to L4 gap, and it is invisible from outside: the codified system looks complete, but the moment a senior leader rotates out, nobody can defend that the change is real.

## L4 - Measured, governed, continuously improved: the org learns at the system level

L4 is rare, and shows up most clearly in isolated pockets - a single product line or delivery team - even when the broader org sits at L2 or L3\. The hallmark is that the org learns at the system level rather than the individual level. Usage, quality, cost, and risk are all instrumented per workflow. These are advanced delivery feedback instruments; the minimum governance, security, value, data, resilience, and workforce gates named in the scope note still apply at every rung and are still reported separately. The instruments belong to this rung, not the one below it: eval sets, the dashboards that carry them, an AI incident taxonomy, and a standing AI ops cadence. A codified delivery system runs without any of them; it just can't tell you whether it worked. When something goes wrong, the response is a system-level adjustment: the prompt library updates, the review checklist gets a new item, the eval suite gains a new failing case.

At L4, the operational risks that haunt earlier rungs are visible as routine telemetry, not crisis discoveries. Drift in model behavior, prompt leakage in production outputs, hallucination at scale, shadow-AI usage outside approved channels. Covered by defined monitoring wherever they are technically observable, with tested detection coverage, an escalation path, and a named owner. Where coverage is real, the team usually sees the signal before a customer complains or a security review escalates. Where it is not, external feedback is still a legitimate detection channel rather than a sign of failure, and the honest move is to name the blind spot instead of assuming the dashboard has none. It is a dashboard line going yellow.

The evidence at L4 is dashboard line items and org-design artifacts, which is what an [honest AI adoption dashboard](https://www.shiftharness.tech/what-an-honest-ai-adoption-dashboard-looks-like/) carries: AI-assisted task ratio per role, cycle-time deltas before and after a workflow redesign, eval-suite pass rate over time, governance-incident counts, approved-tool versus unapproved-tool usage, and named accountability roles for each AI-touched workflow. Every line item maps to a decision someone is empowered to make. Full enablement of a code factory is this rung with decision rights attached: cost and quality tuned deliberately, and learning that crosses project boundaries instead of dying with the team that earned it.

A dashboard without decision rights is not governance. L4 requires a closed loop: signal, owner, decision, action, recheck. If an eval-suite pass rate drops and nobody is empowered to pause a workflow, update a playbook, change a model, or add a gate, the organization is not at L4.

The common failure mode at L4 is the vanity dashboard masquerading as governance. A leadership team commissions an AI dashboard, populates it with the metrics easiest to extract, and presents it monthly with no decision rights attached. That's not L4\. It's L1 or L2 with extra steps, and without the loop the dashboard is theater.

## Self-report drifts upward

The gap between self-reported and observed maturity runs in one direction often enough to expect it: self-report drifts upward. How far is not something this article can tell you, and no inter-rater study is offered here, so measure your own delta rather than inheriting a number. Aggregated across an org, that upward drift is what makes a reported AI maturity curve look healthier than the artifacts justify. A team describing itself as L3 is usually doing solid L2 with a couple of L3 artifacts to point at; a team calling itself L4 is usually sitting on real L3 with a vanity dashboard on top. This is not bad faith. It's the gap between "we have done the work" and "the work has compounded into a system-level capability."

Pull a random sample from the last two sprints: PRs, test plans, requirements documents, ADRs, postmortems, the definition of done, the code review checklist, the prompt library, the AI ops dashboard if one exists. Read them as if you don't work at this company. Place the team at the rung where the artifacts cluster, not the rung where the leadership lives.

A 90-minute artifact review is the screening pass of an AI maturity assessment, not a defensible placement. It produces a provisional placement and a list of unknowns needing corroboration, and it points at the structural gap blocking the next rung, provided three things hold: the sample is random across work types rather than curated, artifact existence is scored separately from active use, and whatever the sample cannot show is recorded as unknown rather than assumed absent. Where those conditions fail, the review tells you what a team documents rather than what a team does, and defensible placement needs a wider sample corroborated against tooling data. The blocking gap is frequently structural rather than technical or cultural, so rule the structural explanation out before accepting either of the others. It is that the next investment on the team's current row (an approved usage baseline, a role-specific workflow change, process integration, or instrumentation) was never built.

The counts in this article (five templates rewritten, three artifacts unchanged, two quarters, 90 minutes) are review prompts drawn from repeated observation, not cutoffs calibrated against independently assessed teams. Use them to structure the conversation, not to settle it. And note what artifact comparison can and cannot show: a difference between this year's artifacts and last year's establishes that the work changed, not that AI changed it. Attribution needs provenance records or tooling telemetry alongside the artifacts.

Use the evidence cluster to place the team:

| Evidence cluster                                                 | Placement |
| ---------------------------------------------------------------- | --------- |
| Artifacts unchanged; no consistent approved use                  | L0        |
| Individuals use AI, but process artifacts unchanged              | L1        |
| Role artifacts improved, but team templates/checklists unchanged | L2        |
| Team process artifacts rewritten and actively used               | L3        |
| Metrics trigger decisions and process updates                    | L4        |

Read the rungs as cumulative: a team sits at the highest rung whose delivery gates are all still satisfied, not at the highest rung it can show one example of. The parallel-dimension gates from the scope note are reported next to the rung, not folded into it. Place the organization where most evidence clusters, not where the best example sits; where a delivery gate is plainly unmet, that caps the rung regardless of where the rest of the evidence sits. One excellent AI-assisted PR does not make a team L2\. One dashboard doesn't make the organization L4\. And define the unit before starting: in an enterprise, run the AI maturity ladder per delivery team rather than once across the whole company, because rungs cluster by product line and a single company-wide number hides every gap worth funding.

![Five delivery artifacts spread on a desk: an annotated PR, a test plan, a requirements doc with tabs, an ADR sketch, and a postmortem, with an abandoned self-assessment sheet set apart.](https://storage.ghost.io/c/73/3e/733efc05-c397-4cf1-b19b-527c1b07dee7/content/images/2026/08/image-2-7.png)

## This article is the journey; the companion piece is the rubric

The ladder places the organization. The rubric, [the 4-level AI adoption evaluation model](https://www.shiftharness.tech/4-level-ai-adoption-evaluation-model/), places the roles inside it. They answer different questions and get used in different conversations.

| Use this article when                     | Use the companion rubric when                          |
| ----------------------------------------- | ------------------------------------------------------ |
| You need to place the organization        | You need to assess a role or person                    |
| You need to decide the next investment    | You need to inspect role-specific artifacts            |
| You need a board-level maturity narrative | You need performance, promotion, or hiring calibration |

The ladder answers "is the org at L3?" The rubric answers "is this senior developer at L3?" Both are necessary, and confusing the two is a common mistake in leadership offsites.

## The diagnostic: five questions you can run before your next leadership offsite

Pull these onto a single page, take a random sample of artifacts from the last two sprints, and answer them honestly. The honest answers produce a provisional placement and a list of unknowns, which is enough to decide the next structural move and not enough to defend a placement to someone who disagrees.

1. **L1 → L2.** Pick a random recent PR, test plan, story, requirements document, and ADR. In each one, can you point to a specific way AI changed how that artifact was produced compared to twelve months ago? If three or more come back "no specific change", the team is no higher than L1 regardless of license count. Distinguishing L0 from L1 needs separate evidence: whether there is any consistent approved use at all.
2. **L2 → L3.** Open the team's definition of done, code review checklist, test plan template, and postmortem template. Were any of them rewritten in the last six months to reference AI-aware steps? If none, the team is still living on individual L2 behaviors without process-level reinforcement.
3. **L3 → L4.** Ask the team's senior engineering manager: if AI was disabled across the toolchain tomorrow, what happens to cadence? "Nothing much" means no higher than L2, because nothing in the process assumed AI; it does not by itself rule out L0 or L1\. Anything else confirms the dependence the process artifacts already implied, and the artifacts are what actually place the team. Record the answer as a resilience reading rather than a rung: "degrades in a way we rehearsed, and here is the cost" is a managed dependency, and "breaks, and we have never tested that" is an unmanaged one that goes in the report beside the rung. Then ask where the dashboard is that tracks AI-assisted task ratio, eval-suite pass rate, and governance incidents. If the answer is "we are building it", the team is at L3 and not yet L4.
4. **L4 stability check.** Take the most recent AI-related incident: a hallucination that reached a user, a model behavior change after a vendor update, shadow-AI usage that came to light. Was it caught through a channel the team's monitoring and escalation design actually covers? If it arrived from outside, was that channel an expected one, and did the response meet the escalation objective? Coverage and response are the L4 signal, not the direction the alert came from.
5. **Self-report vs artifact-evidence delta.** Ask three line managers what rung they think their team is on, then run the artifact review. The gap between the two numbers says more about the org's relationship to evidence than the rung itself.

![Five figures around a meeting table with a page showing five numbered check-marked items; the five-rung ladder framework visible in the background.](https://storage.ghost.io/c/73/3e/733efc05-c397-4cf1-b19b-527c1b07dee7/content/images/2026/08/image-3-7.png)

Most orgs measuring AI adoption today are sitting on the second rung and calling it transformation. Not because the leaders are wrong to want transformation, but because nobody around the table has insisted on placing the org against an artifact-grounded reference model. Once a leadership team has the ladder in front of them and one honest artifact review behind them, the question of what to fund next stops being a debate about tools and becomes a structural question about which asset carries the team off the rung it is actually on. **The next investment is usually not another tool. It is the operating asset the current rung is missing: an approved usage baseline, a role-specific workflow change, process integration, or the instruments to measure what the codified system produces.**

Placing the org against artifact-grounded evidence this way is the lens [Shift Harness](https://www.shiftharness.tech/shift-harness/) applies.

> **AI Transparency Notice:** This article and its accompanying images were created with the assistance of generative AI. The author directed the content, contributed the underlying ideas and analysis, and reviewed the final publication.

## Frequently Asked Questions

What is an AI adoption maturity model?▸

An AI adoption maturity model is a framework that measures how deeply AI has changed how work gets done inside an organization, not how often people open AI tools. This one scores observed behavior in concrete delivery artifacts: pull requests, test plans, requirements documents, architecture decision records, postmortems, definitions of done. The placement reflects what a team produces rather than what it reports. Maturity is the operating-model layer hardening around AI, not procurement.

What are the five levels of AI maturity (L0 → L4)?▸

The five rungs are: **L0 Awareness**, where the team has heard of the tools but nothing in the work product has changed. **L1 Basic tool usage**, where individuals open the tools and paste outputs back in; AI involvement is invisible at the team level. **L2 Role-based workflow usage**, where AI is embedded in how a specific role does its specific work (a PM drafts specs with it, a QA designs test plans with it, a developer reviews diffs with it). **L3 Integrated into the delivery process**, where AI is codified into the applicable parts of the SDLC rather than remembered, with risk-appropriate exclusions documented; removing it would cause a real degradation, and whether that degradation is rehearsed and costed is reported as a separate resilience reading rather than as the rung itself. **L4 Measured, governed, continuously improved**, where usage, quality, cost, and risk are instrumented per workflow; eval suites and governance loops close back into model setup and process design.

How do you assess where a team is on the AI maturity ladder?▸

Pull a random sample of recent artifacts from the last two sprints: pull requests, test plans, requirements documents, ADRs, postmortems, the definition of done, the code review checklist, and the AI ops dashboard if one exists. Read them as if you do not work at this company, and place the team where the artifacts cluster rather than where the leadership lives. A 90-minute review is a screening pass: it produces a provisional range and a list of missing evidence, and names the asset most likely blocking the next rung, provided the sample is random and artifact existence is scored separately from active use. Self-report tends to drift upward, so measure that delta rather than assuming its size.

How is this different from the Gartner or McKinsey AI maturity model?▸

The shapes are similar (most credible models have roughly five levels) but the unit of assessment differs, and the comparison set has changed since this article first published. Gartner's model scores capability across seven abstract categories (strategy, value, organization, people and culture, governance, engineering, data) at five levels, and McKinsey's widely-read State of AI work is survey research rather than a formal appraisal. Two newer entrants do not stop at self-report: the [SEI and Accenture AI Adoption Maturity Model](https://www.sei.cmu.edu/news/sei-and-accenture-release-ai-adoption-maturity-model-to-help-organizations-scale-ai-with-predictable-outcomes/?ref=shiftharness.tech), released in June 2026, spans eight dimensions and was built through practitioner research and Fortune 500 pilots, and CMMI AIM ships formal appraisal and certification assets. Any serious AI maturity model now has to say which evidence it reads. This ladder reads one narrow band deliberately: the delivery artifacts a team already produces, at SDLC altitude. Use an enterprise model when you need cross-dimensional coverage and a formal appraisal; use this ladder when you need a fast, falsifiable read on whether delivery work actually changed. They answer different questions, and neither substitutes for the other.