Skip to content

The Criterion and the Event

Serelora

by Serelora

This article was originally published on Medium.

Read full article on Medium

What a score forgets, and what a gate is for

Photo by Anders Bengs on Unsplash

By Luis Cisneros, CEO of Serelora and COO of EvidenceMD

The missing object

Nobody talks much about the rubric. The conversation drifts, almost automatically, toward the score, the model, the next event, as though the future were a weather system that arrives on its own, indifferent to the rule somebody used when they decided to look. Scores display cleanly. Models have names. The rubric is the line a doctor actually treats as a gate, and it gets handled like a preference buried in a menu. That line is doing the real work.

A chart looks complete when the last field is filled. What you have then is a warehouse of what already happened. The live question starts afterward. Given this kind of patient, and given the rule this kind of doctor was willing to treat as decisive, what happened next.

The sentence is plain enough that it can sound like a footnote. Most of what gets sold as prediction in medicine is an attempt to get around it.

A decoration with a decimal

The ordinary scene is almost dull. A chart is open. Someone invokes a calculator, or an alert fires, or a percentage appears in the sidebar as if it had always lived there. Twenty-two percent mortality. Ten-year ASCVD risk of sixteen percent. Probability of readmission. The figure carries the moral force of a fact because it is precise and numeric. It still has not finished the sentence. Twenty-two percent compared with whom. Drawn from which population. Produced by which calculator, on which date, with missing labs treated as zeros or as absences or as last known values. Against the group this patient actually belongs in, how large is the delta, and is it large enough to change what a competent physician does for this person, in this body, with this remaining time.

Until those pieces are present, the number has the look of knowledge and the substance of decoration.

Someone will say the objection writes itself. Of course the score implies a reference class. That is how scores are built, in the paper, in the methods section, in the population that existed when the authors were still looking at the data. By the time the number reaches the chart, that class has usually evaporated. The percentage remains, unembarrassed, as if the world it was trained on were still standing beside it. An ASCVD estimate of two or four percent in the reference band versus ten or sixteen in this chart is a reason to look. The same raw figure in an eighty-year-old whose baseline already lives outside the textbook interval may not be. Age, frailty, competing mortality, the width of a person’s homeostatic range belong inside the claim. Interpreting that spectrum is the physician’s work. The record has a narrower job. It should keep the comparison class honest and the criterion visible, instead of dropping a naked percent onto a dashboard and calling the result foresight.

The pile and what it forgets

A more ambitious architecture already exists. Gather every longitudinal event into one pile. Train a system to continue the sequence. Ask, for this history, what is likely to happen next. Epic’s Cosmos is the pile. Curiosity is the continuation. The old sepsis score is an ancestor of the same impulse. It compresses a life into a signal that points at an event that has not happened yet. The mountain is the right one. Health systems are not starving for another field. They want the event that failed to arrive. Sepsis intercepted. Readmission avoided. Progression delayed. Death postponed. Insurers have lived inside that question for a long time and gave it a calm name, actuarial science. Medicine arrived later and still keeps mistaking the forecast for the patient.

The pile can be enormous and still forget the useful part.

A next-event model can tell you that something is coming and still hide the rule a person treated as a gate before the something arrived. When the model is right, nobody can say which line to keep. When it is wrong, nobody can say which line to add. You are left with a continuation and a hope that behavior will change downstream.

The rubric in actual use

The unit that will bear analysis is the rubric people actually reach for.

In this sense a rubric is not the guideline in its official clothing. Guidelines are public. They have authors, dates, conflicts, revisions. They can be cited and fought over. The rubric is quieter. It is the one or two lines a given doctor treats as decisive on a Tuesday. Which calculator gets opened. Which threshold is treated as action rather than atmosphere. Which combination of age, lab, comorbidity, and prior event is allowed to move the plan. That is the criterion. The rest is scenery.

Clinics already disclose this if you watch what they reach for. They cluster by specialty because the patients cluster, and the tools follow the patients. Diabetes calculators fire in one shop. Cardiovascular tools fire in another. Metabolic disease groups together because the panel does. The pattern does not require a detective. It requires a record that stops throwing the pattern away.

Named patients make the file worse, not better. They invite heroics, gossip, the romance of the zebra, and a privacy problem that then becomes a reason to do nothing. The interesting object is the repeated gate. This kind of patient. This cutoff. This shop. This later event. De-identify the record. Aggregate at the clinic. Watch how the rubrics move. At that grain the clinic is a phenotype of criteria.

When the rubric is followed and the outcome still arrives

The important case is not the doctor who ignored the rule. It is the shop that carried the rule out word for word and still watched the event happen. That is the moment the question changes. What did the criterion miss.

A snapshot can be obeyed perfectly and still be blind. Fetal growth restriction is an ordinary example. The line many shops treat as the gate is the tenth percentile. Fall below it and the alert fires. Stay above it and the record stays quiet. In one case the measurements never crossed that line. They moved from the fifty-ninth percentile down toward the nineteenth. No single visit met the static cutoff, so nothing notified anyone. The decline itself was the information. A drop of that size across visits is a different kind of risk than a single reading parked just above a published floor. The later outcome still arrived. The rubric had been followed. The longitudinal effect of the trajectory had not been treated as a line in the gate.

Some consensus definitions of late fetal growth restriction already try to name this, a drop of more than two quartiles among the findings that can complete a diagnosis even when a fetus has not sat under the tenth percentile on a given day. The point is not that every shop is using the wrong paper. The point is that the criterion in actual use, the one that generated the silence, weighted the snapshot and did not weight the slope. A reasoning system that only checks whether today’s number cleared yesterday’s threshold will reproduce that silence. A record that keeps the criterion next to the later event can ask a different question. The rule was applied. The outcome still happened. Which line was absent.

Do that across enough offices and the miss stops being an anecdote. The same threshold is carried out. The same kind of patient still gets worse. The gap that keeps appearing is the same gap. That is the material. It is not yet a guideline. It is the reason to run a trial.

The crowd and the shop that moved the outcome

The durable criterion usually sits near the average of what competent specialists already do. Philip Tetlock spent years watching experts forecast. The finding was not that experts are fools. It was that the vivid individual, the person with a theory and a tone of voice, is usually outperformed by the average of careful forecasts. Medicine already runs a quiet version of the same experiment. Thousands of clinicians apply similar tools to similar patients and then live with the next ninety days. That average operating point is a prior. A new cutoff has to earn its keep against it.

The shop whose outcomes move farther than the pack is the exception everyone claims to want and almost no one can see. The record is built to store the score and misplace the rule. The later event lives in another quarter, another system, another billing file, another mortality review that will never be joined back to the threshold that stayed quiet. The outlier appears only when the criterion and the event occupy the same place. Same kind of patient. Two different cutoffs. One shop notified. One shop did not. Ninety days later the event happened, or it did not.

If the threshold stayed quiet and the patient still deteriorated, the rubric is missing a line. The rule failed as a description of the world it was asked to govern. If one shop’s cutoff predicted the slide and the other’s did not, the difference is a result. A defined population. Two named rules. A subsequent event attached.

What counts as a claim

Someone still has to say the result is interesting. An agent can surface the cluster. It can notice that the same miss is appearing across shops that believe they are using the same tool. It can line up the quiet thresholds against the later events. It cannot decide whether the pattern deserves a protocol, a prospective study, or a trial that somebody else already knows how to run. That judgment is about harm, cost, plausibility, and whether the world has already answered the question under another name.

Observational weight can open a narrative and then has to stop. People skip that pause out of impatience dressed up as innovation. Enough cases, enough shops, enough of the same miss, and the slower sequence finally has material. Review. Debate. A randomized trial designed around the missing line. Peer review of that trial. Consensus that the line belongs in the criterion. A guideline that changes what people do on a Tuesday afternoon. That sequence is how a local operating point becomes a public rule. Instrumenting the chart is a way to make the front of it less blind.

The front, at present, is often a hardcoded alert that outlived the paper it was built from. The paper had a population. The alert has a banner. The population dissolved and the banner remained. Alert fatigue gets discussed as if it were a defect in the user. It is what happens when a claim has lost its date, its citation, and its comparison class, and still demands to be obeyed.

Writing the accepted version into every agent first reverses the order. The trial has not happened yet. The consensus does not exist yet. Encoding a suspicion as if it were already a rule is how alerts become fossils. The work is to keep finding the miss, prove the miss, and only then let the public rule change.

The binder

A second institution already works this way and talks as if it were doing something else.

Insurers buy gates. The two dominant gates in American utilization management are InterQual and MCG. One lives under Optum. The other lives under Hearst. Hospitals license them. Plans license them. State Medicaid agencies license them. Medicare Advantage plans may lean on them when Traditional Medicare has not fully established coverage criteria, provided the binder is not used to become more restrictive than Medicare itself. The flavors differ. InterQual tends to match documented findings to granular clinical elements. MCG tends to match a request to an expected pathway and a recovery clock. Both ask whether this patient, at this moment, meets a named definition of medical necessity.

When the patient falls inside the definition, the decision is cheap. A nurse reviewer matches the chart to the line items and moves on. When the patient falls outside it, the cheap procedure ends. Approval does not follow automatically. The expensive procedure begins. The physician has to open a request and do by hand what the record should have been doing. Name the criterion actually in use. Cite the literature that supports it. Argue that this treatment changes the risk of the next event for this kind of person. The conversation is called peer-to-peer review and treated as an administrative ordeal. Conceptually it is a trial of a rubric the licensed binder did not contain.

That ordeal is also a revenue-cycle problem. The binder is the screen through which a claim is allowed to become a payment. If the extra ultrasound, the extra Doppler, the closer interval of surveillance does not match the published line, the plan can refuse it. The physician then spends an afternoon arguing. The plan spends an afternoon defending a margin. Meanwhile the event the extra step was meant to prevent, a missed trajectory, a late diagnosis, a more complex delivery, time in a neonatal intensive care unit, is the payout nobody wanted. Actuarial science is supposed to notice that the cheap denial can purchase the expensive outcome. The fight stays manual because the criterion and the later event are still stored as separate careers.

Binders can be updated. Published guidelines can be updated. The join that almost never gets updated is the one between the criterion and the later event in the same patient class. The reviewer sees whether the note contained the required phrases. The plan sees whether the claim was paid. The clinic sees whether the patient came back worse. Those are supposed to be one fact. They are stored as three careers.

Who keeps the loop in one house

UnitedHealth became difficult to compete with because more of the loop could be kept in one house. The plan holds the risk. Optum holds InterQual, the review staff, the claim edits, pieces of the revenue cycle. Care delivery sits close enough that the encounter is not a rumor. Captive strategy, in this sense, means criterion, event, and payment stop being strangers.

Epic can see a version of the same loop from the other side of the wall. It already sits on the chart. Cosmos already sits on the pile. Curiosity already tries to continue the sequence. Willingness to treat the licensed coverage rule as an object that outcomes can falsify, and then take through a trial, is a different appetite. Information technology operations can instrument a workflow. They do not automatically instrument a claim. Partnerships can close that gap. So can anyone who stores the gate next to what the gate did.

The buyer of loss

CMS is the other buyer, and in some ways the more honest one. Traditional Medicare writes National Coverage Determinations and Local Coverage Determinations. When those are silent, plans fill the silence with internal criteria, often assembled with help from InterQual or MCG. A publicly traded plan extracts margin from the book. The government is trying to reduce loss. Improper payment. Avoidable admission. Care that did not change the event it was purchased to change. Waste that still has to be financed. A state can switch from InterQual to MCG, as Pennsylvania Medicaid did, and describe the change as a contract decision. Underneath the contract is a simpler choice. Which rubric will serve as the screen for necessity.

If de-identified clinic aggregates can show that a named criterion predicts the later event more honestly than the licensed line, the argument is still only observational. The clean way to force the line into public use is the slower one. Run the trial. Survive peer review. Reach consensus. Put the missing line into the guideline. At that point coverage has a problem if it refuses to follow. The plan that withholds the step the new criterion requires is no longer defending a binder. It is accepting the downstream book the trial already priced. A more complex delivery. A neonatal intensive care stay. A preventable deterioration that costs more than the surveillance that would have seen the slope. That is how a criterion becomes something that has to be paid for, without turning every disagreement into another afternoon on the phone.

The government matters here more than any single commercial logo. If the criterion-event loop is tight enough that CMS, a Medicare Administrative Contractor, or a Medicaid agency can use a reviewed rule as a screen and spend less on waste without spending more on harm, the work has left the clinic and entered the country’s accounting system. Governments are not efficient by temperament. They become interested in efficiency when the alternative is to keep paying for events their own rules failed to see coming.

A rule is not a prompt

Falling outside InterQual does not create a right to payment. It creates a demand for a second criterion. The work is to make that second criterion as inspectable as the first. Dated. Cited. Tied to a population. Later tied to an event. A prompt cannot carry that load. A reviewed rule can, after it has been argued in the open. That is the kind of source a reasoning system should be allowed to use. Something that can be shown to a medical director, a utilization nurse, and a regulator without asking any of them to trust a tone of voice.

Two theories of the record

One theory starts with the observatory. Assemble the largest possible pile of events. Train a model to continue the sequence. Hope that a better continuation, shown at the right moment, will change behavior at the bottom. If a new line looks promising, write it into the calculations and wait.

The other starts with the behavior already underway. Doctors are using rubrics, choosing thresholds, ignoring some alerts and obeying others. Reviewers are matching notes to InterQual and MCG. Physicians are arguing on the phone that the binder is missing a line. Sometimes they followed the binder exactly and the patient still got worse. Instrument those choices. Keep the criterion next to the later event. Ask what the rubric failed to see. Let the outcomes argue with the rule. When the same miss appears in enough shops, take it to a trial. Let review and consensus do what they are for. Then the public criterion changes, and the payment screen has to change with it, because refusing the new line is no longer cheap.

AI native is the floor, not the thesis

An AI native record is a cleaner architecture. That part is real. The chart should be addressable in language. An agent should be able to read the file, draft the note, move the order, and keep the physician from typing the chart by hand. We built Serelora that way, including the chat layer, because the old interface wastes time and does not improve the decision.

Chat, agents, ambient capture, and a quieter screen are the right tools for that job. We use them. So will everyone else. Epic will put an agent on Cosmos. Other vendors will put a chat box on the note. Those tools will make documentation faster and alerts easier to read. That is not the same thing as rebuilding what the record is for.

Most of the market is treating those tools as the product. They are features. They make the Tuesday possible. They do not, by themselves, tell you which rule a doctor treated as a gate, or what happened after the rule was applied, or whether the payer should have paid for the step.

What almost nobody is talking about is what the same architecture can do once the criterion and the later event live in one system. A native record can keep the comparison class attached to the number. It can keep an accepted guideline separate from a suspicion that has not been tested. It can store the outcome next to the threshold that was supposed to prevent it. It can turn that pair into data instead of leaving a physician to reconstruct the argument on a peer-to-peer call. That is the work the features are for.

If you only use the agent to predict the next event or to retrieve a citation, you get a faster version of the observatory. The model can point at PubMed and still not know whether the claim should govern the person in the room. A citation tells you where a sentence came from. It does not tell you whether the sentence is true. Of 363 NEJM articles from 2001 to 2010 that tested an existing practice, 146 were reversals. Forty percent. Published is not the same as valid. Prestige is a distribution deal. Cleaning bias out of a measurement does not mean the measurement can explain what happened in the patient.

If you use the same tools to keep the gate and the event together, the record does today’s work and starts producing a kind of data medicine has not had. That is the part missing from the AI native conversation.

What the next record has to emit

The first electronic records did more than move paper onto a screen. They changed what could be counted. Once the chart was no longer trapped in one office, population studies became possible at a scale medicine had not had. Evidence based medicine depended on that change.

That science is now being used as a club. A single paper gets waved at a practicing physician. A model treats the citation as proof. A coverage rule stays in force after the population it was built on is gone. The record that made evidence possible is now helping people confuse a published sentence with a fact about the patient.

The next record has to store something the last generation did not store. The criterion as it was actually used, sitting next to the outcome that followed. The chart still has to handle today’s note, order, and claim. It also has to produce a dataset other clinics can be compared against. That is how a local pattern becomes grounds for a trial. That is how a new line becomes something a payer has to cover. Outcomes and reimbursement improve for the same reason. The rule either changed the event it was paid to change, or it did not.

You do not win by collecting more of the same notes. You win by storing the cheaper object the observatory leaves out.

What we are already building

Serelora is that native record. It structures the data and runs the risk stratification. Chat and agents handle the present work so the gate and the event can sit in one file. That file is what gets handed on.

EvidenceMD does the reasoning over it. Clinical and administrative. It names the gaps against the literature and the next steps that follow. It does not generate the risk. It uses it.

The maternal fetal work is that method in use. Not a parable. A criterion already treated as a gate, joined to the later event, then reasoned over so the miss can be measured. The features run the clinic. The join is the new data those features create once they live in one system.

A clinic without margin cannot do any of this. That is a fact about buildings and payroll. Paying the clinic more does not, by itself, improve the next decision. Survival keeps the lights on. The object of interest stays the same whether the shop is rich or barely standing.

Which criterion, applied to which kind of patient, in whose hands, with what subsequent event. And once money enters the room, whether the person who paid for the treatment bought a change in that event or a ritual that resembled one.

Smaller than an observatory on the first day. Tighter on the question that turns a followed rule and a bad outcome into a public line. Already doing the work of the present, and already generating the stream the next science will require. What the rubric missed, once someone was willing to keep the gate and the event in the same place.