All insights Team HealthMeasurement & ProofThe MethodMoments That MatterRoles & DecisionsFunctions & UnitsIndustriesThe Human Layer & AIOccasionsBehind the Work About Talk to us
Measurement & Proof

"What a Day 14 / 30 / 60 follow-up actually tells you"

Measuring only on the day captures mood, not change. Reading a team again at fourteen, thirty and sixty days separates the improvement that stuck from the glow that faded — and tells a leader what to reinforce.

11 min read

The most common way to measure a team experience is a form at the end of the day. It asks whether people enjoyed it and whether they would recommend it. Those forms come back warm, because people who have just had a good afternoon feel good. Then the number gets filed as proof the thing worked, and everyone moves on satisfied.

It is not proof. It is a photograph of a mood, taken at the one moment mood is highest. The change you actually care about has not happened yet, because change is what is left after the mood is gone. Measuring on the day measures the intervention, not the outcome — and the whole reason to run a team experience is the outcome. This is the argument at the centre of why measurement changes the conversation: the day is the input, and the input is not the result.

Why one reading is not enough

A team experience creates a lift. The useful question is what the lift decays to. Some of it is pure glow and disappears within days — the warmth of a good afternoon, real and pleasant and gone by the following week. Some of it is a genuine shift in how people work together, and it holds. A single reading cannot tell these apart, because on the day they look identical. The team that will be transformed and the team that will be exactly the same by Monday both feel great at 5pm.

So you need to read the team again, after the glow has had time to fade, and then again after normal work has had time to test whatever remains. This is not measurement as a verdict handed down once. It is measurement as a series of readings that, together, reveal the shape of what actually happened. That is why we read at three points after an experience, not one — and why each of those points is chosen to answer a specific question the others cannot.

What each window is for

Fourteen days. This is the first honest reading. The afternoon is over, the group chat about it has gone quiet, and people are back in the ordinary run of work. Whatever shows up now survived the glow. If a shift in how openly a team talks, or how it makes a decision, is still visible at fourteen days, it is real rather than residual. If everything has snapped back to exactly where it was, you have learned something important and inexpensive: the day was enjoyable and nothing changed. Fourteen days is the earliest point at which the answer is trustworthy, because it is the earliest point at which the mood is no longer distorting the reading.

Thirty days. By now the team has hit normal pressure — a deadline, a disagreement, a bad week, the ordinary friction of work. Thirty days tells you whether the change holds when things get hard, which is the only condition under which it matters. Improvements that evaporate the first time work gets stressful were never really improvements; they were good intentions that could not survive contact with a difficult Tuesday. A change still standing at thirty days has been tested by real conditions and passed. That is the difference between a team that felt more trusting in a calm room and a team that actually behaves more trustingly when the stakes rise — and only the second one is worth anything.

Sixty days. This is the question of whether the change has become the new normal or is quietly slipping back. At sixty days a genuine shift has either settled into how the team simply works now — invisible, unremarked, just how things are done here — or it has faded and the team has drifted back to its old baseline without anyone quite noticing. Sixty days is long enough that habit has either formed or failed to. Knowing which is what lets a leader decide whether to reinforce, repeat, or move on to a different gap. A change that has become the baseline needs protecting, not repeating. One that is slipping needs a different response than one that never took at all.

We are not going to set out here what each reading consists of, because that belongs to the work and not to this page. The point is the cadence itself, and what the shape across the three windows reveals.

Why these intervals, and not others

The specific spacing is not arbitrary, and it is worth understanding why. Read too early — a day or two after — and you are still measuring the glow; the reading tells you how good the afternoon felt, not what changed. Read only once, far out, and you miss the shape entirely: you cannot tell a change that held steady from one that spiked and decayed and happened to be caught on the way down. The three windows are spaced to catch three different things — survival of the glow, survival of pressure, and settling into habit — because those are three distinct tests a change has to pass, and they happen at different times.

There is nothing sacred about the exact days; what matters is the logic behind them. Soon enough to catch the fade, far enough to catch the test, far enough again to catch the settling. A team's improvement is not a single event you can photograph once. It is a trajectory, and you need at least three points to see the shape of a curve. One point is a dot. Two can be joined by any line you like. Three begins to tell the truth about direction.

Reading the shape

Put three readings next to the starting point and a team's response takes on a shape, and the shape is far more informative than any single value. A sharp rise that falls away by fourteen days was a good day and nothing more — enjoyable, forgettable, no change to reinforce. A smaller rise that holds through thirty and sixty days is a real change that will compound, because it has already proven it can survive both the fade and the pressure. A dimension that barely moves at first and then climbs across the sixty days is often a slow-burning shift in trust, which rarely announces itself immediately and instead accumulates as the team gathers evidence that the new way is safe.

And a curve that rises, holds at thirty, then slips at sixty tells its own story: something real happened but was not reinforced, and the team is drifting back. That is a signal to act, and act specifically — not to conclude the experience failed, but to see that it worked and then faded for want of reinforcement. Those shapes are useful in a way a single number never is. They tell a leader not just whether something worked, but what kind of thing happened, and therefore what to do next. A number says pass or fail. A shape says what to do on Monday.

Two teams, two curves

Make it concrete with two teams that had, on the day, identical experiences and identical enthusiasm. Team A comes off a well-designed session on committing to decisions. At fourteen days the reading shows decision-making up and holding. At thirty days the team has been through a genuine disagreement and the improvement survived it — they committed, and stayed committed, under real friction. At sixty days it has become simply how they operate; no one talks about the offsite anymore because the change has stopped being a change and become a habit. That is the shape of a success, and the sixty-day flatness at the higher level is the best outcome there is: the new normal.

Team B had the same day, the same warmth, the same high feedback score. At fourteen days their reading is up too. But at thirty days, the first hard week, decision-making has slipped back toward baseline, and by sixty days it is exactly where it started. Same day, same enthusiasm, opposite result. The end-of-day form recorded both as successes and would have sent both teams back to the same catalogue next year. The cadence tells you Team A's design worked and should be built on, and Team B's produced warmth without change and needs something different — probably a look at why the change could not survive pressure, which usually points a dimension deeper, often to trust. Without the follow-ups, that entire distinction is invisible, and both teams get treated as wins.

Different dimensions move differently

One reason the shape matters is that the eight dimensions of team health do not all move on the same schedule. Some respond fast and fade fast; some barely move at first and then build. Energy can lift immediately and also drain back quickly, so a change in it shows early and needs watching. Trust is the classic slow burn — it is built from accumulated evidence, as we describe in trust inside a team, so a genuine trust shift often reads as flat at fourteen days and only climbs once the team has had enough safe instances to update its default. Alignment can jump quickly if the experience produced real clarity, then hold or erode depending on whether the clarity gets reinforced in normal work.

Reading only once would flatten all of this into a single misleading snapshot. A trust improvement checked only at fourteen days would read as a failure, when it was actually just early. The cadence, read across the eight dimensions, is what lets you see each dimension on its own timeline, which is the only way to tell a slow success from a real failure. That distinction alone justifies the whole practice, because getting it wrong means abandoning a change that was about to take, or reinforcing one that already faded.

You cannot read a curve without a starting point

None of this works without a reading from before the experience. The follow-ups measure change, and change is a comparison; with nothing to compare against, a reading at thirty days is just a standalone snapshot that tells you where the team is, not how far it moved. The before-reading — the scan that starts our method — is what turns the follow-ups into a measured trajectory rather than a series of disconnected observations. It sets the baseline the curve is drawn from.

This is why measurement and diagnosis are two ends of one discipline, not separate services. The scan names the target and sets the zero point; the follow-up cadence tracks movement against it. Take away the scan and the follow-ups lose their meaning. Take away the follow-ups and the scan is a diagnosis no one ever checked. Together they make a claim possible: this dimension was here, we aimed at it, and here is the curve showing where it went.

What a leader does with it

The cadence turns measurement from a verdict into a guide. It shows which changes are sticking and which are slipping, so reinforcement goes where it is needed rather than everywhere or nowhere. A leader who knows that the team's decision-making improved and held, but its cross-team trust spiked and faded, knows precisely what to reinforce and what to leave alone — which is a far more useful position than "the offsite was a success" or "the offsite was a waste."

It gives the board a curve instead of a claim, which matters enormously for whether the programme survives a budget review — the case we make for the People leader in proving the ROI of culture spend and for the executive in the retention math. And over several experiences it builds a picture of how a particular team responds — how fast its energy fades, how slowly its trust builds — which makes every subsequent design sharper. The cadence is not just measurement of one event. It is the accumulating knowledge of a team, which is the most valuable thing a People function can hold.

What reinforcement actually looks like

Reading the curve is only useful if a leader does something with it, so it is worth being concrete about what reinforcement means when a change is holding but fragile. It is rarely another offsite. More often it is small and continuous: a leader who noticed at thirty days that the team's new openness was starting to slip can deliberately protect it — keep making it safe to raise hard things, keep responding well when someone does, keep modelling the behaviour the experience introduced. The experience plants the change; reinforcement is the daily watering that decides whether it takes root or dies back.

The cadence tells a leader exactly where to spend that effort, which is its practical payoff. Without it, reinforcement is either absent — the team is left to drift back on its own — or scattered evenly across everything, which is the same as nowhere. With it, the leader knows the decision-making change is holding on its own and needs no attention, while the cross-team trust is slipping and needs deliberate protection this month. That precision is the difference between a change that compounds and one that quietly reverses while everyone assumes the offsite handled it. A great many team experiences that "did not work" actually worked and were then allowed to fade, unreinforced, because no one was reading closely enough to catch the slip while it was still catchable.

Why most measurement stops at the day

If the later readings are so much more informative, why does almost everyone stop at the end-of-day form? Partly effort: the form is easy and the follow-up takes discipline and a system. Partly timing: by fourteen days everyone has moved on to the next thing, and no one wants to reopen a box already ticked. And partly, honestly, comfort — the end-of-day form reliably returns a warm number, and the follow-up risks returning an inconvenient one. It is more pleasant to file the happy score than to learn, three weeks later, that the happy score decayed to nothing.

But the comfort is exactly the problem. Measurement that only ever confirms is not measurement; it is reassurance. The value of the cadence is precisely that it can tell you a day did not work, which is the only kind of information that improves what you do next. A company that measures only on the day can run team experiences for years and never learn a thing about which of them changed anything, because it never once looked at the moment the truth becomes visible.

A cadence, not a one-off

The follow-up works because it is a rhythm, not a special event someone remembers to run. The moment it depends on goodwill and memory, it fails — by fourteen days everyone is busy, by thirty the box feels closed, and the reading that never happens tells you nothing. So the readings have to be built into the calendar from the start, scheduled the moment the experience is booked, treated as part of the work rather than an optional extra someone might get to. A follow-up that is nice-to-have is a follow-up that will not happen.

This is also what makes the cadence compound over time. A team read on this rhythm across several experiences accumulates a history — how fast its energy fades, how slowly its trust builds, which kinds of design hold and which evaporate. That history is worth more than any single reading, because it turns each new design from a guess into an informed choice built on how this specific team has actually responded before. The cadence is not just how you prove one experience worked. It is how a People function slowly comes to genuinely know the teams it serves, which is the foundation everything else is built on. Run once, it is a check. Run as a habit, it becomes knowledge.

The lightest version that still works

Some leaders hear "read the team four times" and picture a burden the team will resent. It does not have to be heavy. The readings are proportionate — light enough that the team barely registers them, deliberately so, because a follow-up that feels like an audit changes the very behaviour it is trying to observe. The discipline that matters is not volume but consistency: the same dimensions, read the same way, at spaced intervals, taken honestly. A modest reading done every time beats an elaborate one done once and abandoned.

The overhead people fear is almost always the overhead of doing it badly — turning it into a heavy survey, a big analysis, a report no one reads. Done well, the cadence is quiet and small on the team's side and disciplined on ours. The weight is in the consistency, not the size of each touch. A leader worried about burden should worry far more about the far larger burden of spending on teams for years with no idea what any of it did.

The proof is in the weeks after

The day is the intervention. The fortnight, the month, and the two months after it are the proof. A team experience that is never read again is an experience whose result no one will ever know — it might have transformed the team or changed nothing, and without the later readings those two outcomes are indistinguishable and will stay that way forever. Skipping the follow-up is how a company spends on its teams for years and never learns whether any of it worked, repeating whatever felt good and dropping whatever did not, guided entirely by mood.

Read the team before, to set the baseline. Read it at fourteen days, to see what survived the glow. Read it at thirty, to see what survived pressure. Read it at sixty, to see what became the new normal. The curve those readings draw is the difference between hoping a team experience worked and knowing what it did. That knowledge is the entire point of measuring, and it lives in the weeks after the day, not in the day itself. A company that learns to read the weeks after stops guessing about its teams and starts steering them, which is the whole reason to measure at all.

Common questions

Why measure a team at 14, 30 and 60 days instead of just once?

Because a single reading captures mood on the day. Fourteen days shows whether anything survived the initial glow, thirty days whether it held under normal work pressure, and sixty days whether it has become a new baseline or faded.

What does the follow-up tell a leader to do?

It shows which changes are sticking and which are slipping, so the leader knows exactly what to reinforce rather than guessing or assuming the day fixed everything.

Isn't a same-day feedback score enough?

No. A same-day score measures enjoyment at the moment enjoyment is highest, which is the least informative moment. The change worth paying for is what remains after the glow fades, which only later readings capture.

Do you need a reading from before the experience too?

Yes. Without a baseline reading there is nothing for the follow-ups to be compared against. The before-reading is what turns the later ones into a measured change rather than a standalone snapshot.