A caregiving AI pilot, the admin panel

December is two working weeks long, so the analysis is built before the data arrives.

The study ends on 4 December and the analysis runs to 18 December, which is two working weeks with a holiday closure at the end of them. Turning nine weeks of responses into answers is real work, and it has to happen either way.

Doing that work now, against invented data, is what makes December the part where we read the answers rather than the part where we build the thing that produces them.

How it was built

Each thing could only be made once the thing before it existed.

Not quite the order the calendar says, because all of it moved at once. It is the order the work depends on, which is the part that matters.

Start

The research plan

What the study is trying to learn, the measures it will be judged on, the groups of people each one is counted among, and the calendar it all runs on.

From the plan

The instruments

The four surveys, the interview guides, the diary prompts and the contact schedule. You cannot write a question until you know which measure it has to answer, so nothing here could be drafted before the plan settled.

From four things at once
The instrumentswhat is actually asked The admin panel we already hadso nothing needed is lost The research planwhat counts, and on whom Design rules of my ownhow a number may appear

This panel

The instruments decide what can be shown at all, because a panel cannot display something nothing collects. The existing panel decides what must not be lost. The plan decides what counts and against which group. And the design rules decide how it appears: counts before percentages, every figure naming its population, nothing not measured drawn as a zero, and no view that cannot say who acts on it.

And one arrow pointing back

Building the screen sent a requirement back the other way. Four fields on the participant card turned out to be collected by no instrument in the study, which was only visible once somebody tried to draw them. The panel found a gap in the instruments rather than the other way round.

How it is used

It is not one panel. It is three jobs on three different clocks.

Every view names who acts on it and what they do. A view that cannot name its own consequence is not a view, and two features were cut for failing that test.

What it runs on

Product activity arrives on its own. Everything else is a spreadsheet somebody uploads, so those figures are only ever as current as the last upload. Each file has its own slot, and the panel shows what changed before anything is saved.

Product activityarrives on its own
The participant trackerresearch firm and Care Guide · weekly
Recruitment screenerrecruitment partner · one-off
Survey responsesstudy lead · each wave
Deploy logback end developer
Clock oneEvery week, from Week 1
What happens

Somebody passes 21 days without using any of the three main features. Logging in and looking at the dashboard do not count and do not reset the clock.

What the panel shows

The inactive list, with who has not been contacted yet, marked as this week's work.

Who acts

The study lead, or whoever is running the week.

What they do

Offer that person a short interview about why they stopped. The target is two or three of them, not all of them.

The panel says on its own screen that this is a working queue rather than a health check. It also names its own blind spot: attending a Care Guide session counts as activity, so if the tracker’s attendance columns have not been ticked yet, somebody may be on this list who should not be.
Clock twoFour times, as each survey wave closes
What happens

A wave closes. Intake, Week 2, Week 6, then exit.

What the panel shows

How many people answered, and how that compares with the waves before it.

Who acts

The study lead.

What they do

Chase the people who have not answered personally, rather than sending another reminder, and leave the window open longer.

This is the clock with real stakes. Seven of the nine measures live in the exit survey alone, and there is no second chance to collect them.
Clock threeOnce, from 7 December
What happens

The exit survey closes on 7 December and the last upload lands.

What the panel shows

All nine measures at once, each with its target, its reading, the group it was measured on, and its verdict.

Who acts

The team, in the readout.

What they do

Read the answers. Not build the thing that reads them.

Nothing settles before December. No measure carries a verdict until the exit wave closes, which is why the panel spends most of its life showing what it is waiting for rather than a score.

Drawn from what the panel states on its own screens, not from a description of it. The activity, the dates and the people are invented so the panel could be reviewed before the study opens.

Where the numbers come from

Every figure traces to a named file and the day a person uploaded it.

There are no live connections, so the panel never says the word connected. It says who uploaded what, and when, and it goes stale loudly rather than quietly.

SourceWhoHow oftenWhat it carries, and what breaks without it
Product activityNobodyContinuousThe only feed that arrives on its own. Messages, forecasts opened, days used.
The participant trackerTwo update it, the study lead uploads itWeeklyThe sheet the research firm keeps, and the single most load-bearing feed here. It carries who is enrolled, who is in the depth group, whether each Care Guide session was actually attended, and each person’s contact status. Three people keep it true: the research firm updates enrolment and contact status, the Human Care Guide ticks off each session after it happens, and the study lead exports it and uploads it here each week. Every percentage on the panel divides by it, and it is also the exclusion list, so without it the figures are not merely stale, they silently include test accounts. It is also where the two identifiers meet. The research firm issues a participant identifier and the participant has a username in the product, and both sit on the tracker, which is what lets a survey answer be joined to what somebody actually did. That join used to be a file somebody built by hand, and it was the largest single source of data that could not be matched up. It reaches the panel as a weekly export from one worksheet built for the purpose, which carries the study columns and no names, no emails and no payment amounts. The working sheet keeps everything personal and is never exported.
Recruitment screenerResearch firmOnce, before Week 1How much care somebody provides and whether they live with the person they care for. It is what decides who is worth interviewing when somebody goes quiet.
Survey responsesStudy leadFour timesExported straight out of the form tool without renaming a column. Seven of the nine measures live in the last of these four.
Deploy logBack end developerAfter each releaseWhich build a person was answering about. A survey wave that straddles a release averages two different products into one figure, which is exactly how an earlier study lost most of its metrics.

Every upload is kept, so a wrong file can be rolled back to the one before it. Before anything saves, the panel shows how many rows moved, any column it did not recognise, and anybody in the file who is not on the roster.

A recreation

The scoring screen, rebuilt small enough to put on a page.

This is the part of the panel that decides whether the thing worked. Nine measures, each with a target agreed before anybody was recruited, and each naming the group of people it is measured on. Nothing here is typed in. Every percentage below is worked out in your browser from the counts underneath it, the same way the real screen works it out from the uploaded file. The dashed line on each bar is the target.

Switch it between the two states. Mid-study is the honest picture six weeks in: the exit survey has not been sent, so seven of the nine cannot be scored at all and the screen says so rather than showing a zero.

Material in, answers out

It takes in raw material and gives back answers, not more material.

The product knows a great deal. The point of the panel is that almost none of that reaches a screen as it arrived, because a screenful of raw material is somebody else's work to do, not a finished thought.

What goes in

Everything people typed to the assistant, as running text
Every action taken in the product, timestamped
Survey answers from four waves, one response at a time
Errors, and what was actually said when they happened
Screener answers about how much care each person gives
Which build somebody was using when they answered

What comes out

Who has gone quiet, and who is worth talking to this week
How many people answered each wave, and how that changed as the study went on
Each measure against its target, on a named group of people
Which errors could have harmed somebody, and who would catch them
Where every figure came from, and the day a person uploaded it
What the study could not measure, and why not

Getting it back out again is a feature, not a courtesy.

The last study had no export. Everything anybody needed came out by screenshotting a screen, or by copy and pasting while scrolling and hoping the selection held. That cost somewhere around ten hours, and none of those hours produced anything: it was the same numbers, moved by hand, into a document that then had to be checked against the screen it came from.

So this one has an export button on every view that shows a figure, and it was in the plan before any of the views were. A panel that can only be read is a panel you have to transcribe, and transcription is where numbers go wrong.

Four things are left out on purpose, and the panel says so on its own screen rather than leaving somebody to notice. What people typed to the assistant, because a conversation carries health information about people who are not in the study and never agreed to be in it. What the AI costs to run, because the study is not asking about money. Any inactivity cut other than 21 days, because 21 is the one trigger anybody agreed to act on and offering 14 or 30 invites a number nobody will do anything with. And account management, which is the only thing here that changes the product rather than reading from it, so it is kept apart from every page that only reads.

The first of those is the one worth saying out loud. It is not a feature that got lost, it is the research plan being enforced on a screen.

The change worth arguing about

The safety screen changed, we lost something doing it, and we put it back.

The panel we already have counts boundaries held and boundaries missed as two totals, side by side. This one sorts errors by whether they could have harmed somebody medically, emotionally or financially, which is what the team asked for in August, and it treats one as enough to stop.

Those are not the same information, and for a while the second one quietly replaced the first. On the older screen 273 missed sits beside 230 held, and 273 alone reads very differently from 273 beside 230. A boundary that was correctly held had become invisible.

It is back. Each person's record now shows their safety events one at a time, each marked as a boundary held or a boundary missed, because a boundary the product correctly held is not a failure and collapsing the two into one number is what makes a safety figure unreadable.

What is still not there is the pair as cohort totals, the way the older summary put 273 next to 230. That is the remaining question rather than a decision already taken, and it is an addition rather than a rebuild.

What I tried

Four things were built and then taken out, and the removals are the design.

Rebuilt

The first version had its numbers written into the page

It looked finished. It was a picture of a panel. The figures had been transcribed carefully and still disagreed with the data they came from, which is the whole argument against transcribing. It now works every figure out from the answers behind it.

Cut

A watch list at 14 days, with an alert on it

Nothing in this study re-engages a participant who goes quiet, so the alert had no owner and no next step.

Cut

A queue for safety events, running from untriaged to resolved

A workflow with states implies a team working through it and an end it reaches. There is exactly one decision available here, and the tiers that queue needed do not exist yet.

Narrowed

Every number opens into its detail

Inherited from an earlier specification and pulled back. Most figures already show their detail in the same view, so opening was ceremony. Two things open now, a person's name and a measure's reading, because they are the only two that hide anything.

The numbers mean nothing to anyone outside the few people who wrote them, and a document that says “KR 1.3 cannot be scored” forces every reader to go and look it up, or worse, to nod along without knowing what was lost.Why every measure on this page is named rather than numbered
What I learned

Four things I would tell somebody starting the same job.

Agreeing a definition is not agreeing a screen
We settled the measures in the summer and still disagreed about the panel in August, on things nobody knew were open. The picture surfaced them, not the reading.
A figure typed by hand is a promise you will not keep
The first version's numbers were transcribed carefully and were already wrong against their own source. If the page cannot work a number out, that number drifts the first time anything changes.
If a view cannot name who acts on it, it is decoration
This test removed two features that both felt responsible and diligent, and it is the most useful question to ask of anything on a dashboard.
Build the empty state first
The panel spends its first six weeks mostly waiting, and that is the version most people will see.

The panel is not the study. It is December's work, done in August.