
TL;DR: Progress iteratively with feedback is the third ITIL 4 guiding principle. Do not try to do everything at once. Break the work into small pieces, ship one, and use feedback to steer the next. Generative AI makes this cheaper than it has ever been, you can ship an AI change in an afternoon, which is exactly why the discipline matters more, not less. Two failure modes dominate: the big-bang AI program that runs for a year and proves nothing, and the deploy-and-forget agent nobody grades. The test is simple: can you ship one small AI change, measure its real effect, and let that number decide whether to widen or correct? Pilot on one team, grade every result, expand only on evidence. With GenAI the model is a commodity. The feedback loop is the product.
Part of the ITIL 4 series: The seven ITIL 4 guiding principles, The ITIL 4 service value chain, rebuilt with AI Coworkers, Focus on value, the first guiding principle, Start where you are, the second guiding principle.
This is the third post in a series on the seven ITIL 4 guiding principles. Part 1 was Focus on Value, what to aim at. Part 2 was Start Where You Are, where to aim from. This one is about how to move: in small steps, with a feedback loop wired in from the start. Nowhere does that matter more than with AI.
ITIL 4 says to organize work into small, manageable iterations, each delivering something of value, with feedback sought and used during and between them. It is a caution against two things at once: analysis paralysis, where you plan forever and ship nothing, and the grand rollout, where you ship everything at once and learn nothing until it is too late to change cheaply.
The word that carries the weight is feedback. An iteration without a feedback loop is just a smaller version of the same gamble. The principle is not "ship small." It is "ship small, then let what you learn change what you do next." Feedback is designed in, not bolted on at the end.
There are two ways teams break this principle, and generative AI pushes on both.
The first is the big-bang AI program. A twelve-month "AI transformation" that lands as one enormous rollout. By the time it ships, the assumptions it was scoped on are stale, nobody remembers why half of it was built, and there is no clean way to tell which part actually helped. It felt fast because it was busy. It failed slow.
The second is newer and sneakier. Generative AI made shipping almost free, so teams push an AI answer, a deflection, an agent live in an afternoon, and then never grade it. This is deploy-and-forget, and it is worse than it looks. An AI change you do not measure is not progress. It is a liability with a good demo. The chatbot that confidently gives the wrong answer does not announce itself. You find out from the second ticket the user files after they gave up on the first.
GenAI did not repeal this principle. It raised the stakes on it.
Here is the whole principle in one move. Pick one request type. Ship an AI change for just that, an answer, a deflection, a drafted resolution a human approves. Then measure the effect that actually matters, not "it responded," but did the person get unblocked, did the ticket stay closed, was the answer correct. Then do one of two things: widen it, or correct it.
Notice what the measurement is not. It is not "we deployed AI." It is not tickets touched. It is the outcome: resolved on first contact, no reopen, no silent second attempt. If you cannot state the one number that means it worked before you ship, you are not iterating. You are hoping.
We built a read-only report meant to show the value IT was delivering. The first version measured a usage proxy and called it value. We showed it to the person who would actually use it, and the feedback was blunt: that is not real value. The tempting response was to add more to it. We did the opposite. We shipped a smaller, sharper version that measured one thing well, showed it again, took the next round of feedback, and only then widened it to the full picture. Three tight loops beat one big build, and the only reason we could turn them that fast was that generative AI made each rewrite cheap.
Same story with the service desk analyzer. We shipped the core analysis, put it in front of a real need, and a real gap came back: a team rolling out a new tool with no test environment had nowhere to start. That feedback, not a roadmap, drove the next iteration.
To make it concrete, here is that loop as a single AI Coworker. You point it at one deflected category, say password and access how-to. It does not answer tickets. Each week it reads that category read-only and grades whether the AI answers actually resolved the request: no reopen, no re-ask by the same person, no escalation to a human, no thumbs down. It puts a true-deflection rate on the board, shows the trend against last week, and, most usefully, lists the exact answers that failed and the knowledge article each one used.
Then it does the disciplined thing. If the score is below the bar, it recommends holding: fix those articles, re-measure. Only when the score clears the bar does it recommend widening to the next category, and a human makes that call. It never edits the knowledge and never widens on its own.
When the model is a commodity, and it increasingly is, the durable advantage is not which model you picked. It is the loop you built around it: how you capture whether it worked, how you grade it, how fast you correct it. If you are going to ship AI fast, build the grading in on day one: a way to see the deflection that failed, the answer that was wrong, the automation that ran when it should have asked.
Start where the feedback will be honest and the blast radius is small. One team. One queue. People who will tell you the truth when it breaks, because they have to live with it. Let them find the failure modes while they are cheap. Then expand where the evidence supports it, not where the launch plan says you should be by now.
Pick one high-volume request type. Ship one small AI change for it, and if you want it safe, ship a draft-only version where the AI proposes and a human approves. Write down the one number that would mean it worked, and the before number, so you can prove the difference. Measure for a week. Then correct or widen, based on what the number says, not on how the demo felt.
What does "progress iteratively with feedback" mean in ITIL 4? It is the third of the seven guiding principles. It means organizing work into small iterations that each deliver value, and actively seeking and using feedback during and between them, rather than delivering everything in one large effort.
Why does this principle matter more with generative AI? Because GenAI makes shipping trivial and measuring optional. It is now easy to deploy an AI change in an afternoon and never check whether it helped. The discipline of a feedback loop is what turns a convincing demo into a system you can trust.
How small should the first AI iteration be? One request type, one team, one measurable outcome. Small enough that you can ship it this week and read the result next week.
Isn't iterating slower than a single big rollout? No. Big rollouts feel fast because they are busy, then fail slowly and expensively. Small loops surface the problems while they are still cheap to fix.
Is this an official ITIL resource? No. ITIL 4 is owned by AXELOS and PeopleCert. This post explains the guiding principle in our own words; it is not affiliated with or endorsed by AXELOS or PeopleCert.
Generative AI did not make this principle optional. It made it the whole game. When anyone can ship an AI feature in an afternoon, the advantage moves to whoever can measure honestly and correct fastest. Ship small. Grade every result. Widen on evidence, not on optimism. The teams that win with AI will not be the ones with the biggest rollout. They will be the ones with the tightest loop.
Next in this series: Collaborate and Promote Visibility. For the overview of all seven, see the guiding principles.
Sources and further reading: PeopleCert ITIL 4 Foundation; AXELOS ITIL service management.