Tokenising the patient journey: from records to representations
— by Vivax
A Lancet Perspective asks how time-aware patient representations could turn fragmented records into computable clinical context—while keeping the boundary…
A chart is not a patient journey. It is a set of observations made at different moments, in different formats and for different purposes: a diagnosis code, a laboratory result, a note, an image, a procedure and a timestamp. Mahmood and Topol’s Lancet Perspective argues that temporal AI foundation models could take that ordered, multimodal history and produce a computable representation. Here, tokenising means expressing clinical events and data in a form a model can work with; it does not simply mean de-identifying a record.
The proposed shift is from isolated documents and one-off tasks to a representation that can retain sequence and context. Codes, note text, imaging, laboratory trends and procedures may each contribute evidence, but their meaning changes with what came before and after. A time-aware representation is therefore an engineering abstraction of an evolving record, not a complete digital twin of a person. Representation learning can organise clinical context; it does not by itself establish cause, benefit or safe action.
It is important to read the source at its proper weight. “Tokenising the patient journey: from records to representations” is a one-page Perspective by Faisal Mahmood and Eric J Topol in The Lancet, published 5 September 2026. It reports no new cohort, model training, benchmark or clinical endpoint. Its value is a thesis and a map to relevant literature—not evidence that its authors built, validated or deployed a clinical system.
One cited example is Zhang and colleagues’ Apollo preprint. The authors report 25 billion records from 7.2 million patients across 28 modalities and 12 specialties, with a held-out test set of 1.4 million patients. They describe 322 prognosis and retrieval tasks, including 95 new-disease-risk tasks, and forecasts of risk up to five years. Those are reported preprint benchmark results, not clinical proof or a result of the Lancet Perspective. The scale makes the evaluation question concrete; peer review, prospective utility and transferability still require separate evidence.
A second cited line of work, Delphi-2M by Shmatko and colleagues, models sequences of disease states. Its Nature report describes training with UK Biobank data, validation in Danish registries and 1,258 states or tokens; it reports rates for more than 1,000 diseases and sampled future trajectories or disease-burden estimates up to 20 years. A long forecast horizon is not a treatment recommendation. Such outputs do not establish causality, fairness, calibration across settings, privacy protections, prospective usefulness or who remains accountable for care.
For Vivax, the practical question is not whether a fluent system can narrate a chart, but whether a representation stays grounded in records, time and clinical evaluation. Vivax’s position is that a clinical world model must be evaluated as a record-grounded, time-aware representation under clinical governance—not judged by fluent text alone. We are building toward unified representations of patients, procedures and clinical events over time, while keeping claims bounded by the evidence: no Apollo-equivalent model, validated virtual patient, predictive performance or autonomous clinical capability is claimed here.