commvita
Connected care platform
Platform & architecture

Data architecture and data quality

Three models kept apart, a dictionary that can say no, a rule that every number drills to its source, and quality owned by the service that makes the data. What’s measured, what’s seeded, and what’s still thin.

Live vs demonstrated: Live — real, API-backed platform logic (wired end-to-end today) Demonstrated — representative control surface with seeded data / illustrative UI mock-up
© 2026 Commvita Digital Health Solutions Ltd. All rights reserved.

1 One record, three models, never folded together

Most data problems in health aren’t about storage. They’re about two systems using the same word for different things, or the same thing under different names, and nobody noticing until a figure on a board pack is wrong. commvita’s data architecture starts from that failure and is built to make it visible.

The platform holds three models and keeps them apart on purpose. The canonical model is NHS England’s: a subset of the national canonical data model, pinned and checked field by field against the publisher, with no drift found at the last check. The semantic model is commvita’s own: a dictionary that records what each column across the estate means and whether two meanings are the same. The physical model is the tables underneath both. The Data Models screen shows all three and never presents one as if it were another, because folding them together is how a reader comes to believe a platform implements an object it merely names.

The Data Models screen showing the pinned canonical model: sixteen objects, 507 fields, 111 link types and an entity diagram with Patient at the centre
The national canonical model as commvita pins it: sixteen objects and the eighteen relationships it can resolveCaptured from the running system, build B-692 · demonstration data

Two habits on that screen carry into everything else. A relationship is only drawn when the platform can resolve it. Where the national model declares a link from both ends and gives a backing field on neither, the link is reported as unchecked instead of inferred from a column that looks like a key. And an object with no relationship inside the pinned set is ringed and left alone. Guessing a join from a name is the single quickest route to a confident wrong answer, and the platform doesn’t take it.

2 What a column means, and who says so

The semantic spine is where the estate’s columns are given meanings. Not by a rule that matches names, but by a person recording a judgement, one binding at a time, with the false matches written down beside the true ones.

The semantic spine coverage tab: 704 models, 8,810 columns, eleven concepts, 1,141 columns bound, all drafted and none signed off, 12.95 per cent coverage, and 7,669 columns not yet bound
The honest state of the dictionary: a little over an eighth of the estate bound, and nothing yet signedCaptured from the running system, build B-692 · demonstration data

The figures are unflattering and that’s the point of publishing them. Of 8,810 columns across 704 models, 1,141 are bound to one of eleven concepts, and every one of those bindings is a draft. Sign-off is a named person accepting accountability for a meaning, and the platform never sets it. The screen says in words that commvita doesn’t own these meanings and that this is commvita’s own estate, not the national single patient record’s spine.

The judgement that matters most is not comparable. A column called patient_name is a label, and two people can share it. A column called patient_id on most tables is a module-local reference that resolves to nobody in the register. Both are recorded as not comparable to the person reference, so no query can join on them and no report can count them as people. A registry that can only say “equivalent” asserts comparability by omission. This one can say no.

The semantic model tab with eleven concepts and a graph where a dashed edge records that two meanings were judged not comparable
Six relationships recorded between eleven concepts, most of them a person saying two things are not the sameCaptured from the running system, build B-692 · demonstration data

3 Where a number comes from

Every count, badge and picklist on a screen is a projection of a canonical dataset, and a projection is only legitimate if it can be traced back to the records that produced it and behaves the way its source implies. That rule is a standing order across the platform, and it’s why the tiles on this site’s screenshots drill to people.

The standard has plain consequences. A count you can’t click into is treated as a defect. A dropdown is bound to the one canonical dataset for its entity, by stable identifier and never by label, and every option has to resolve to something the next module along can read with the same key. When the source has no matching records the control says so, and never shows a cached option that no longer exists upstream. And the person comes first: almost every clinical interaction begins with someone in front of a professional, so the record is anchored on the person and the numbers are anchored on the record.

The Architecture screen carries a data-lineage classification for the modules that matter most at go-live: which fields staff type, which the system generates, and which are seeded on the demonstration estate and would arrive by integration in a deployment. It is the document a migration lead reads first.

The data lineage and entry classification table: 30 modules classed as manual entry, 11 as system-generated and 3 as seeded, each listing its manually entered and system-generated fields
Which fields people type and which the platform makes, module by moduleCaptured from the running system, build B-692 · demonstration data

4 Quality is measured, then owned by the service

Data quality in the NHS fails when it belongs to an information team and to nobody else. The platform’s model is the one national data-quality programmes have converged on: one agreed definition per metric, ownership by the service that creates the data, measurement instead of assumption, issues that are visible and have to be acted on, and improvement done where the data is made.

The Data Quality Management overview with four headline measures, the five service-led ownership principles, and a per-module table of completeness, accuracy and timeliness scores
The five principles, and a score per module for completeness, accuracy and timelinessCaptured from the running system, build B-692 · demonstration data

The rule library is the part worth reading slowly. Seventeen checks of the kind an acute or community trust runs against its own patient administration data: appointments left as booked after the event, missing ethnicity, discharge outcomes that disagree with an open pathway, duplicate admissions, group names where a clinician’s name should be, pre-migration referrals that never moved. Each carries the impact in plain words, the module it belongs to, the number of affected records, a severity and a status. Issues raised from the rules sit on their own tab with an owner and a detection date.

The data quality rule library listing twelve of seventeen checks with description, impact, module, affected records, severity and status
Twelve of the seventeen rules, each saying what it catches and why that mattersCaptured from the running system, build B-692 · demonstration data
Be clear about which figures are measured. The Data Quality Management screen is a demonstrated surface: its headline percentages and per-module scores are seeded, because the endpoints it asks for aren’t wired on this build, and the provenance register classes the page as mixed. The measured data-quality work lives elsewhere. The research extract runs a data-quality dashboard against the real cohort and returns review until a vocabulary is loaded. The platform’s own data-quality summary counts, live, the patients with no NHS number or date of birth, the referrals with no source and the open record conflicts. And the semantic spine’s coverage figure above is measured from the schema on every request.
The research data-quality dashboard run on a real cohort, returning a review verdict with completeness, conformance and plausibility checks and per-table row counts
The measured one: a data-quality run on a live cohort that says review, and names the reasonsCaptured from the running system, build B-692 · demonstration data

5 Time and identity, the two hard columns

Two things go wrong in every large clinical estate: when something happened, and who it happened to. The platform measures both about itself and publishes the answer.

Time. Of the estate’s timestamp columns, 444 still hold a time as text. That figure is measured from the live schema by the platform’s own position endpoint, and a build check refuses any change that makes it worse. The event spine was built before the clean-up, with its one occurrence column born typed and mandatory, so the record of what happened doesn’t inherit the debt the clean-up exists to pay. Each store then gets a typed twin, dual-read and reversible, and nothing is deleted.

Identity. Sixty-two endpoints resolve a person through the same function. It normalises identifiers to digits, resolves an identifier held by more than one person to nobody instead of the first match, and a row that can’t be bound to a person says so in words instead of showing the key. The register measured 10,545 staff-naming references that resolve to no account and 560 columns that name the patient under ten different spellings. Those are the numbers the seam exists to make harmless.

Reconciliation is a first-class job. The platform’s test suite has thirty files whose only purpose is a concept that was found split across two stores that couldn’t see each other: deprivation of liberty, the SNOMED crosswalk, appraisals, immunisations, messages, billing. Each suite asserts that the legacy store folds in on read, that the response shape is unchanged, and that the reconciliation key is normalised, because matching raw strings reconciles nothing while reporting success.

6 What it isn’t, and where it’s thin

It isn’t a data warehouse. Nothing here is a copy kept for reporting; every screen reads the record the people doing the work write to, and research extracts are cut per cohort on demand with a ledger entry. It isn’t a graph database, and standing one up wouldn’t make a graph: the relationships are computed in one place at read time, and the right first move was to record events, which is what was built.

The honest edges. The semantic spine binds a little over an eighth of the estate and nothing is signed off. The Data Quality Management module is seeded end to end on this build. No crosswalk yet joins the canonical and semantic models, and nothing claims one. The national canonical pin is a logical conformance check against a published specification, never a comparison with a running national instance. And 444 columns still hold time as text.

Where it lives in commvita

ScreenRouteWhat it showsStatus
Data models/data-modelsCanonical pin, semantic registry, physical schema● Live
Semantic spine/semantic-spineConcepts, bindings, coverage measured from the schema● Live
Architecture · Data/architectureLineage and entry classification per module○ Demonstrated
Data Quality Management/data-qualityPrinciples, rule library, issues, actions○ Demonstrated
Research data-quality run/omop-cdmCompleteness, conformance and plausibility on a live cohort● Live
Ontology position/ontology/positionTimestamp and event-producer position measured from the live schema (API)● Live
© 2026 Commvita Digital Health Solutions Ltd. All rights reserved. NHS FDP canonical data modelFHIR R4 · SNOMED CT · dm+dOMOP CDM v5.4Non-SaMD